0

Magentic Marketplace Eval

Fresh

Magentic Marketplace (multi-agent agentic-markets benchmark) as a verifiers v1 eval environment

Type
RL Env
License
mit
Size
v0.3.0
Published
Sep 2026
Updated
Sep 2026

Cite

Notes

Only stored in your browser.

magentic-marketplace-eval

Magentic Marketplace (multi-agent agentic-markets benchmark) as a verifiers v1 Environments-Hub eval environment. Vendored from https://github.com/microsoft/multi-agent-marketplace (arXiv 2510.25779, MIT).

What the model under evaluation plays

  • Every LLM marketplace seat: all customers AND all businesses — one seat per agent, one verifiers interaction per seat. Customers run the benchmark's autonomous-shopping loop (search / message / check / pay / end), businesses respond to inquiries with the benchmark's inquiry-response prompt (text or structured order proposals).
  • Each upstream generate() decision call becomes one turn on that seat's interaction; structured-output parse retries become nudge turns on the same interaction (mirroring upstream's append-the-error-and-retry loop).
  • An episode = one full marketplace simulation: the in-process FastAPI marketplace server + SQLite store + all agent loops run on a worker thread; seat turns are bridged onto the episode's event loop with asyncio.run_coroutine_threadsafe and served concurrently under the episode's agent gate.

Scenario cells (benchmark-experiment fidelity)

Each cell = bundled population x marketplace search algorithm (the benchmark's own market-design axis) x bounded customer step budget:

  • mexican3_simple / mexican3_lexical / mexican3_optimal — the paper's quickstart population (mexican_3_9: 3 customers, 9 businesses) under rating-ranked / lexical-relevance / fully-fulfilling search. 12 seats.
  • mexican5_lexical — first 5 customers / 15 businesses of mexican_10_30, lexical search; the market-size axis (3 -> 5 customers). 20 seats.
  • contractors5_simple / contractors5_lexical — first 5 customers / 15 businesses of contractors_10_30 (home-services domain); cross-domain generalization cells. 20 seats.
  • mexican3_simple_tight / mexican5_lexical_tight / contractors5_lexical_tight — the same three populations with a 6-decision customer budget instead of 10. The standard budget saturates at these market sizes (per-cell means 0.91-1.00), so the tight budget is the discrimination axis; nothing else changes between the pairs.

9 cells x 2 seeds = 18 tasks (the rollouts per task supply the sampling variance; the seed axis labels tasks, since the simulation has no RNG upstream of LLM decisions).

Upstream's 40-seat populations are STRUCTURALLY UNRUNNABLE here, not merely contended: one 40-seat episode can consume the org-wide ~32-request model ceiling by itself (measured: 39% of seats lost to provider 429s), and in PrimeAgentHarness mode it also exhausts the team's 2000-tunnel quota. They stay bundled but are not a cell.

Cell size is deliberately bounded at 12-20 seats. Upstream's full bundled populations (10 customers / 30 businesses = 40 seats) were measured to be unviable: each seat's agent run consumes a concurrency slot (the platform's org-wide ~32-request ceiling is shared by every running eval) and, in PrimeAgentHarness mode, a team tunnel slot — 40-seat cells died with Maximum number of team tunnels (2000) reached while 12/20-seat cells ran cleanly. Medium cells deterministically subset a bundled population (first N in file order); the full populations remain bundled and reachable through the taskset knobs. The env's default seat gate is max_concurrent_agents: 8 for the same reason.

Direct-model seats default to a LOCAL SUBPROCESS runtime (SubprocessConfig): the null-harness chat program is tool-less, so it runs on the runner VM with no container and no interception tunnel. Prime-runtime seats mint one tunnel per rollout and the team's tunnel quota (2000) is shared across every running eval, so staying local removes that failure class. Pin "agent": {"runtime": {"type": "prime", "cpu": 8, "memory": 8}} for PrimeAgentHarness-mode seat runs (those need a container).

Cells x seeds = 12 tasks. Rewards recorded per seat trace: mm (composite, weight 1) + mm_<component> (weight 0) with raw quantities in trace.info["magentic"].

Metrics (documented components, all in [0, 1], higher = better)

  • purchase_completion — fraction of customers who paid for at least one proposal (upstream purchase_completion_rate).
  • needs_met — fraction of customers whose paid orders matched the full requested menu AND amenities (upstream needs-met flag).
  • welfare — total customer utility (upstream utility: 2x requested menu value - payments for needs-met purchases) normalized by the population's theoretical optimum (cheapest fully-matching business per customer).
  • proposal_validity — fraction of order proposals without integrity errors (off-menu items, wrong menu prices, inconsistent totals, nonexistent agents; upstream proposal-error classes).

The composite is the unweighted mean. All raw quantities + per-customer / per-business records land in trace.info["magentic"]["raw"].

Vendoring + deviation notes (magentic_marketplace)

  • magentic_marketplace/ is vendored top-level from packages/magentic-marketplace/src/magentic_marketplace (MIT, LICENSE.magentic-marketplace in this repo).
  • DB: upstream runs on PostgreSQL (docker-first). This port runs on the upstream SQLite controller (platform/database/sqlite/) — episodes need no docker/postgres and state stays inside the episode sandbox. Analytics run over the live SQLite controller with the upstream MarketplaceAnalytics engine.
  • LLM layer: upstream dispatches to OpenAI/Anthropic/Gemini structured outputs. This port installs a seat bridge at the vendored marketplace/llm entrypoint (functional.install_seat_bridge): calls route to the verifiers seat interaction. Structured outputs are emulated on the plain-chat seat by appending the JSON schema to the decision prompt (upstream's prompts rely on provider-side structured outputs, so the schema suffix is an eval-specific addition) and parsing + bounded retry (3 attempts, mirroring upstream).
  • Pruned (unused on the eval path): the FastAPI UI static assets, the postgres-only CLI/experiment tooling (cli.py, experiments/list_experiments.py, experiments/export_experiment.py, experiments/run_audit.py). The postgres connector remains but imports lazily (asyncpg not required).
  • Determinism: no RNG upstream of LLM decisions (search ranking, proposal validation, analytics are deterministic given actions); the seed axis labels tasks; the stochastic element is model sampling.
  • Empty-reply guard (copied from the agent-bazaar-eval template): a bounded self-heal nudge turns empty model replies into a retry, otherwise the episode fails loudly (SeatPoisonedError); counters recorded into trace.info["magentic_seat"].

Running

Local:

.venv/bin/vf-eval magentic_marketplace_eval --env-dir-path . \
    -m internal/glm-5.3-fast -n 12 -r 1 --max-concurrent 1 \
    --max-tokens 131072 --env.agent.runtime.type subprocess \
    --no-serve --no-rich

Hosted (canary first, then the full 12):

prime eval run primeintellect/magentic-marketplace-eval --hosted \
    -m internal/glm-5.3-fast -n 1 -r 1 --max-concurrent 1 \
    --max-tokens 131072 --timeout-minutes 60 \
    --eval-name mm-canary --plain

PrimeAgentHarness seat mode (probes): add --env-args '{"agent": {"harness": {"id": "prime-agent"}}}' and for seat-heavy cells pin "runtime": {"cpu": 8, "memory": 8} + "max_concurrent_agents": 4.