econarena-eval
EconArena economic-games benchmark as a verifiers v1 Environments-Hub eval environment.
PAPER-FAITHFUL REIMPLEMENTATION of the alpha games described in "Economics Arena for Large Language Models" (Guo et al., arXiv:2401.01735, 2024). No official code exists (verified: the paper and every citing paper reference only the arXiv preprint; no repo exists on GitHub/Gitee/HuggingFace) — the games, session structure, answer formats and prompts are reconstructed from the paper's Sections 3-4 and Appendices A (games/metrics/architecture) and B (verbatim prompt templates). Prompt-fidelity is therefore approximate where the paper used placeholders; this package documents every reconstructed parameter.
What the model under evaluation plays
The model plays every contestant seat in a session (the paper's
"self-competing" configuration): 5 contestants, one verifiers interaction
each, the paper's per-run decision prompts, runs runs per session
(default 10; the paper's history experiments use 6+ runs per session with up
to 3 runs of history revealed).
- beauty_L / beauty_M / beauty_H — beauty contest games (Keynes' "guess 2/3 of the average"): pick a real number in [0, c̄]; closest to 2/3 of the average wins a fixed prize (ties split); unique NE = 0. The paper's range groups: c̄ drawn from [10,100) / [100,1000) / [1000,10000). No game history revealed (the paper's baseline environment).
- beauty_hist — L-range beauty contest with revealed history (the paper's B.4 prompt): after run 1 each prompt carries the last 3 runs of everyone's choices. This is the paper's strategic-reasoning / in-context-learning manipulation; convergence over runs is the signal.
- auction_L — minimal second-price sealed-bid (Vickrey) auction cell, the paper's group L (private values s ~ N(50,10), assets 100, entrance fee 10 on rule breaks, minimal-id tie-break). Truthful bidding is the unique symmetric NE. Auctions are otherwise deprioritized in this program — auctions/bargaining are covered by the GTBench port; this one cell is kept only because the paper's second game family is cheap to include.
Metrics (paper Section 3.2 / Appendix A.3, as documented components)
Components in [0,1], higher = more rational; composite EAS (EconArena
Score) = unweighted mean, recorded at reward weight 1 with components at
weight 0 and all raw quantities + per-run series in trace.info["econarena"]:
rationality— the paper's Eq. 1: mean payoff / NE payoff (beauty: prize share vs prize/5 under the all-tie NE; auction: assets vs the run's NE asset).ne_proximity— 1 − mean normalized NE-deviation distance. NOTE: the paper's ratio formulad = π(a)/π(NE) − 1is undefined at the beauty-contest NE (π(0) = 0), so the deviation is implemented as the paper's deviation distance: |action − NE action| / scale (c̄ for beauty, assets for auctions).win_rate— beauty: mean prize share won; auction: fraction of runs won.rule_following— fraction of runs with a valid in-range parsable action (the paper's rule-breaking frequency, inverted).convergence— NE proximity over the LAST THIRD of the runs (the paper's "convergence rate to the optimal NE strategies"; the in-context learning signal on history cells).
Reconstructed parameters (not fixed by the paper)
Fixed prize x = 100 (beauty); entrance fee = assets/10 (auction); 5 contestants; 10 runs per session default; history window = 3 runs; sampling temperature 0.7 (the paper does not state one). Everything is exposed in task info so episodes are reproducible.
Running
Local (model-free stub tests first):
pytest tests/
Local live smoke (2-run session, 1 task, chat program runs locally):
.venv/bin/eval econarena_eval -m internal/glm-5.3-fast -n 1 -r 1 \
--max-concurrent 1 --max-tokens 131072 --no-serve --no-rich \
--env.agent.runtime.type subprocess \
--env-args '{"taskset": {"cells": ["beauty_L"], "seeds": [8], "runs": 2}}'
Hosted (canary probes only; 1 episode, 2-run and 12-run variants):
prime eval run primeintellect/econarena-eval --hosted \
-m internal/glm-5.3-fast -n 1 -r 1 --max-concurrent 1 \
--max-tokens 131072 --timeout-minutes 60 \
--eval-name econarena-t2-direct --plain \
--env-args '{"taskset": {"cells": ["beauty_L"], "seeds": [8], "runs": 2}}'
PrimeAgentHarness mode: add "agent": {"harness": {"id": "prime-agent"}} to
--env-args (seats then run as Prime Agent daemons in a container runtime).
Provenance and licensing
This package contains no vendored upstream code (none exists). The engine
(econarena_eval/engine.py) is an original reimplementation of the paper's
rules and prompts, MIT-licensed here. Cite the benchmark as: Guo, Bu, Wang,
Ren, Sui, Shang, Lu. "Economics Arena for Large Language Models."
arXiv:2401.01735 (2024).