magentic-marketplace-eval
Magentic Marketplace (multi-agent agentic-markets benchmark) as a verifiers v1 Environments-Hub eval environment. Vendored from https://github.com/microsoft/multi-agent-marketplace (arXiv 2510.25779, MIT).
What the model under evaluation plays
- Every LLM marketplace seat: all customers AND all businesses — one seat per agent, one verifiers interaction per seat. Customers run the benchmark's autonomous-shopping loop (search / message / check / pay / end), businesses respond to inquiries with the benchmark's inquiry-response prompt (text or structured order proposals).
- Each upstream
generate()decision call becomes one turn on that seat's interaction; structured-output parse retries become nudge turns on the same interaction (mirroring upstream's append-the-error-and-retry loop). - An episode = one full marketplace simulation: the in-process FastAPI
marketplace server + SQLite store + all agent loops run on a worker
thread; seat turns are bridged onto the episode's event loop with
asyncio.run_coroutine_threadsafeand served concurrently under the episode's agent gate.
Scenario cells (benchmark-experiment fidelity)
Each cell = bundled population x marketplace search algorithm (the benchmark's own market-design axis) x bounded customer step budget:
mexican3_simple/mexican3_lexical/mexican3_optimal— the paper's quickstart population (mexican_3_9: 3 customers, 9 businesses) under rating-ranked / lexical-relevance / fully-fulfilling search. 12 seats.mexican5_lexical— first 5 customers / 15 businesses of mexican_10_30, lexical search; the market-size axis (3 -> 5 customers). 20 seats.contractors5_simple/contractors5_lexical— first 5 customers / 15 businesses of contractors_10_30 (home-services domain); cross-domain generalization cells. 20 seats.mexican3_simple_tight/mexican5_lexical_tight/contractors5_lexical_tight— the same three populations with a 6-decision customer budget instead of 10. The standard budget saturates at these market sizes (per-cell means 0.91-1.00), so the tight budget is the discrimination axis; nothing else changes between the pairs.
9 cells x 2 seeds = 18 tasks (the rollouts per task supply the sampling variance; the seed axis labels tasks, since the simulation has no RNG upstream of LLM decisions).
Upstream's 40-seat populations are STRUCTURALLY UNRUNNABLE here, not merely contended: one 40-seat episode can consume the org-wide ~32-request model ceiling by itself (measured: 39% of seats lost to provider 429s), and in PrimeAgentHarness mode it also exhausts the team's 2000-tunnel quota. They stay bundled but are not a cell.
Cell size is deliberately bounded at 12-20 seats. Upstream's full bundled
populations (10 customers / 30 businesses = 40 seats) were measured to be
unviable: each seat's agent run consumes a concurrency slot (the platform's
org-wide ~32-request ceiling is shared by every running eval) and, in
PrimeAgentHarness mode, a team tunnel slot — 40-seat cells died with
Maximum number of team tunnels (2000) reached while 12/20-seat cells ran
cleanly. Medium cells deterministically subset a bundled population (first N
in file order); the full populations remain bundled and reachable through the
taskset knobs. The env's default seat gate is max_concurrent_agents: 8 for
the same reason.
Direct-model seats default to a LOCAL SUBPROCESS runtime
(SubprocessConfig): the null-harness chat program is tool-less, so it runs on
the runner VM with no container and no interception tunnel. Prime-runtime seats
mint one tunnel per rollout and the team's tunnel quota (2000) is shared across
every running eval, so staying local removes that failure class. Pin
"agent": {"runtime": {"type": "prime", "cpu": 8, "memory": 8}} for
PrimeAgentHarness-mode seat runs (those need a container).
Cells x seeds = 12 tasks. Rewards recorded per seat trace: mm (composite,
weight 1) + mm_<component> (weight 0) with raw quantities in
trace.info["magentic"].
Metrics (documented components, all in [0, 1], higher = better)
purchase_completion— fraction of customers who paid for at least one proposal (upstreampurchase_completion_rate).needs_met— fraction of customers whose paid orders matched the full requested menu AND amenities (upstream needs-met flag).welfare— total customer utility (upstream utility: 2x requested menu value - payments for needs-met purchases) normalized by the population's theoretical optimum (cheapest fully-matching business per customer).proposal_validity— fraction of order proposals without integrity errors (off-menu items, wrong menu prices, inconsistent totals, nonexistent agents; upstream proposal-error classes).
The composite is the unweighted mean. All raw quantities + per-customer /
per-business records land in trace.info["magentic"]["raw"].
Vendoring + deviation notes (magentic_marketplace)
magentic_marketplace/is vendored top-level frompackages/magentic-marketplace/src/magentic_marketplace(MIT,LICENSE.magentic-marketplacein this repo).- DB: upstream runs on PostgreSQL (docker-first). This port runs on the
upstream SQLite controller (
platform/database/sqlite/) — episodes need no docker/postgres and state stays inside the episode sandbox. Analytics run over the live SQLite controller with the upstreamMarketplaceAnalyticsengine. - LLM layer: upstream dispatches to OpenAI/Anthropic/Gemini structured
outputs. This port installs a seat bridge at the vendored
marketplace/llmentrypoint (functional.install_seat_bridge): calls route to the verifiers seat interaction. Structured outputs are emulated on the plain-chat seat by appending the JSON schema to the decision prompt (upstream's prompts rely on provider-side structured outputs, so the schema suffix is an eval-specific addition) and parsing + bounded retry (3 attempts, mirroring upstream). - Pruned (unused on the eval path): the FastAPI UI static assets, the
postgres-only CLI/experiment tooling (
cli.py,experiments/list_experiments.py,experiments/export_experiment.py,experiments/run_audit.py). The postgres connector remains but imports lazily (asyncpg not required). - Determinism: no RNG upstream of LLM decisions (search ranking, proposal validation, analytics are deterministic given actions); the seed axis labels tasks; the stochastic element is model sampling.
- Empty-reply guard (copied from the agent-bazaar-eval template): a
bounded self-heal nudge turns empty model replies into a retry, otherwise
the episode fails loudly (
SeatPoisonedError); counters recorded intotrace.info["magentic_seat"].
Running
Local:
.venv/bin/vf-eval magentic_marketplace_eval --env-dir-path . \
-m internal/glm-5.3-fast -n 12 -r 1 --max-concurrent 1 \
--max-tokens 131072 --env.agent.runtime.type subprocess \
--no-serve --no-rich
Hosted (canary first, then the full 12):
prime eval run primeintellect/magentic-marketplace-eval --hosted \
-m internal/glm-5.3-fast -n 1 -r 1 --max-concurrent 1 \
--max-tokens 131072 --timeout-minutes 60 \
--eval-name mm-canary --plain
PrimeAgentHarness seat mode (probes): add
--env-args '{"agent": {"harness": {"id": "prime-agent"}}}' and for
seat-heavy cells pin "runtime": {"cpu": 8, "memory": 8} +
"max_concurrent_agents": 4.