agent-bazaar-eval
Agent-Bazaar (THE_CRASH + LEMON_MARKET) as a verifiers v1 Environments-Hub eval environment. Vendored from https://github.com/CameronCrow/Agent-Bazaar (MIT).
What the model under evaluation plays
- crash_baseline / crash_stabilizing: all 5 B2C firms — one seat per firm, one verifiers interaction per seat, the benchmark's own per-step pricing prompts (supply / production / price in one JSON decision per step).
- lemon_main / lemon_norep: all 12 C2C buyers — one seat per buyer; the benchmark's bid prompts (bid/pass on visible listings) and post-purchase review prompts (upvote/downvote). Sellers and the Sybil cluster are fed from the bundled listing corpus, so only buyers make model calls.
Each task = one full market episode for one (scenario cell, seed); 4 cells x
3 seeds = 12 tasks. Rewards recorded per seat trace: eas (composite,
weight 1) + eas_<component> (weight 0), with raw quantities and per-step
series in trace.info["bazaar"].
Scenario cells (benchmark-experiment fidelity)
- crash cells replicate
scripts/exp1.py_BASE_FIXED(THE_CRASH, 5 LLM firms, 50 CES consumers, dlc=3, io prompts, no diaries), with--num-stabilizing-firms 0|3. - lemon cells replicate
scripts/exp2.py_BASE_FIXED(LEMON_MARKET, 12 buyers, K=3 sybils, rho_min 0.3, rep 0.8, dlc=3), plus the bundled listing corpus;lemon_norepadds--no-buyer-rep. - 50 timesteps per episode by default (crash dynamics emerge within that horizon; lemon keeps the benchmark default).
Reproducibility
Seeds are applied exactly as agent_bazaar.main does. The benchmark draws
from process-global RNGs during stepping, so run episodes one-per-process
for exact reproducibility: --serve.max-concurrent 1, or --max-concurrent 1
in-process. Throughput then scales with pool workers (processes).
Running
Local:
prime eval run agent_bazaar_eval --env-dir-path . -m internal/glm-5.3-fast \
-n 12 -r 1 --max-concurrent 2 --max-tokens 131072
Hosted (canary first, then the full 12):
prime eval run primeintellect/agent-bazaar-eval --hosted \
-m internal/glm-5.3-fast -n 1 -r 1 --max-concurrent 1 \
--max-tokens 131072 --timeout-minutes 60 --eval-name bazaar-canary --plain
Vendoring notes (agent_bazaar)
agent_bazaar/is vendored top-level (absoluteagent_bazaar.*imports).- Patched:
agents/llm_agent.pyimports the provider model clients lazily;models/__init__.pyships onlyBaseLLMModel;main.pystrips wandb and is used only forcreate_argument_parser(). agent_bazaar_eval/corpus/listing_corpus.jsonis the repo's pre-compiled listing corpus (15.5 MB) so lemon episodes need no seller LLM calls.