0

Agent Bazaar Eval

Fresh

Agent-Bazaar economic-alignment benchmark (THE_CRASH + LEMON_MARKET) as a verifiers v1 eval environment

Type
RL Env
License
mit
Size
v0.1.3
Published
Sep 2026
Updated
Sep 2026

Cite

Notes

Only stored in your browser.

agent-bazaar-eval

Agent-Bazaar (THE_CRASH + LEMON_MARKET) as a verifiers v1 Environments-Hub eval environment. Vendored from https://github.com/CameronCrow/Agent-Bazaar (MIT).

What the model under evaluation plays

  • crash_baseline / crash_stabilizing: all 5 B2C firms — one seat per firm, one verifiers interaction per seat, the benchmark's own per-step pricing prompts (supply / production / price in one JSON decision per step).
  • lemon_main / lemon_norep: all 12 C2C buyers — one seat per buyer; the benchmark's bid prompts (bid/pass on visible listings) and post-purchase review prompts (upvote/downvote). Sellers and the Sybil cluster are fed from the bundled listing corpus, so only buyers make model calls.

Each task = one full market episode for one (scenario cell, seed); 4 cells x 3 seeds = 12 tasks. Rewards recorded per seat trace: eas (composite, weight 1) + eas_<component> (weight 0), with raw quantities and per-step series in trace.info["bazaar"].

Scenario cells (benchmark-experiment fidelity)

  • crash cells replicate scripts/exp1.py _BASE_FIXED (THE_CRASH, 5 LLM firms, 50 CES consumers, dlc=3, io prompts, no diaries), with --num-stabilizing-firms 0|3.
  • lemon cells replicate scripts/exp2.py _BASE_FIXED (LEMON_MARKET, 12 buyers, K=3 sybils, rho_min 0.3, rep 0.8, dlc=3), plus the bundled listing corpus; lemon_norep adds --no-buyer-rep.
  • 50 timesteps per episode by default (crash dynamics emerge within that horizon; lemon keeps the benchmark default).

Reproducibility

Seeds are applied exactly as agent_bazaar.main does. The benchmark draws from process-global RNGs during stepping, so run episodes one-per-process for exact reproducibility: --serve.max-concurrent 1, or --max-concurrent 1 in-process. Throughput then scales with pool workers (processes).

Running

Local:

prime eval run agent_bazaar_eval --env-dir-path . -m internal/glm-5.3-fast \
    -n 12 -r 1 --max-concurrent 2 --max-tokens 131072

Hosted (canary first, then the full 12):

prime eval run primeintellect/agent-bazaar-eval --hosted \
    -m internal/glm-5.3-fast -n 1 -r 1 --max-concurrent 1 \
    --max-tokens 131072 --timeout-minutes 60 --eval-name bazaar-canary --plain

Vendoring notes (agent_bazaar)

  • agent_bazaar/ is vendored top-level (absolute agent_bazaar.* imports).
  • Patched: agents/llm_agent.py imports the provider model clients lazily; models/__init__.py ships only BaseLLMModel; main.py strips wandb and is used only for create_argument_parser().
  • agent_bazaar_eval/corpus/listing_corpus.json is the repo's pre-compiled listing corpus (15.5 MB) so lemon episodes need no seller LLM calls.