0

Econarena Eval

Fresh

EconArena economic-games benchmark (beauty contest + second-price auction) as a verifiers v1 eval environment — paper-faithful reimplementation of ...

Type
RL Env
License
mit
Size
v0.1.2
Published
Sep 2026
Updated
Sep 2026

Cite

Notes

Only stored in your browser.

econarena-eval

EconArena economic-games benchmark as a verifiers v1 Environments-Hub eval environment.

PAPER-FAITHFUL REIMPLEMENTATION of the alpha games described in "Economics Arena for Large Language Models" (Guo et al., arXiv:2401.01735, 2024). No official code exists (verified: the paper and every citing paper reference only the arXiv preprint; no repo exists on GitHub/Gitee/HuggingFace) — the games, session structure, answer formats and prompts are reconstructed from the paper's Sections 3-4 and Appendices A (games/metrics/architecture) and B (verbatim prompt templates). Prompt-fidelity is therefore approximate where the paper used placeholders; this package documents every reconstructed parameter.

What the model under evaluation plays

The model plays every contestant seat in a session (the paper's "self-competing" configuration): 5 contestants, one verifiers interaction each, the paper's per-run decision prompts, runs runs per session (default 10; the paper's history experiments use 6+ runs per session with up to 3 runs of history revealed).

  • beauty_L / beauty_M / beauty_H — beauty contest games (Keynes' "guess 2/3 of the average"): pick a real number in [0, c̄]; closest to 2/3 of the average wins a fixed prize (ties split); unique NE = 0. The paper's range groups: c̄ drawn from [10,100) / [100,1000) / [1000,10000). No game history revealed (the paper's baseline environment).
  • beauty_hist — L-range beauty contest with revealed history (the paper's B.4 prompt): after run 1 each prompt carries the last 3 runs of everyone's choices. This is the paper's strategic-reasoning / in-context-learning manipulation; convergence over runs is the signal.
  • auction_L — minimal second-price sealed-bid (Vickrey) auction cell, the paper's group L (private values s ~ N(50,10), assets 100, entrance fee 10 on rule breaks, minimal-id tie-break). Truthful bidding is the unique symmetric NE. Auctions are otherwise deprioritized in this program — auctions/bargaining are covered by the GTBench port; this one cell is kept only because the paper's second game family is cheap to include.

Metrics (paper Section 3.2 / Appendix A.3, as documented components)

Components in [0,1], higher = more rational; composite EAS (EconArena Score) = unweighted mean, recorded at reward weight 1 with components at weight 0 and all raw quantities + per-run series in trace.info["econarena"]:

  • rationality — the paper's Eq. 1: mean payoff / NE payoff (beauty: prize share vs prize/5 under the all-tie NE; auction: assets vs the run's NE asset).
  • ne_proximity — 1 − mean normalized NE-deviation distance. NOTE: the paper's ratio formula d = π(a)/π(NE) − 1 is undefined at the beauty-contest NE (π(0) = 0), so the deviation is implemented as the paper's deviation distance: |action − NE action| / scale (c̄ for beauty, assets for auctions).
  • win_rate — beauty: mean prize share won; auction: fraction of runs won.
  • rule_following — fraction of runs with a valid in-range parsable action (the paper's rule-breaking frequency, inverted).
  • convergence — NE proximity over the LAST THIRD of the runs (the paper's "convergence rate to the optimal NE strategies"; the in-context learning signal on history cells).

Reconstructed parameters (not fixed by the paper)

Fixed prize x = 100 (beauty); entrance fee = assets/10 (auction); 5 contestants; 10 runs per session default; history window = 3 runs; sampling temperature 0.7 (the paper does not state one). Everything is exposed in task info so episodes are reproducible.

Running

Local (model-free stub tests first):

pytest tests/

Local live smoke (2-run session, 1 task, chat program runs locally):

.venv/bin/eval econarena_eval -m internal/glm-5.3-fast -n 1 -r 1 \
    --max-concurrent 1 --max-tokens 131072 --no-serve --no-rich \
    --env.agent.runtime.type subprocess \
    --env-args '{"taskset": {"cells": ["beauty_L"], "seeds": [8], "runs": 2}}'

Hosted (canary probes only; 1 episode, 2-run and 12-run variants):

prime eval run primeintellect/econarena-eval --hosted \
    -m internal/glm-5.3-fast -n 1 -r 1 --max-concurrent 1 \
    --max-tokens 131072 --timeout-minutes 60 \
    --eval-name econarena-t2-direct --plain \
    --env-args '{"taskset": {"cells": ["beauty_L"], "seeds": [8], "runs": 2}}'

PrimeAgentHarness mode: add "agent": {"harness": {"id": "prime-agent"}} to --env-args (seats then run as Prime Agent daemons in a container runtime).

Provenance and licensing

This package contains no vendored upstream code (none exists). The engine (econarena_eval/engine.py) is an original reimplementation of the paper's rules and prompts, MIT-licensed here. Cite the benchmark as: Guo, Bu, Wang, Ren, Sui, Shang, Lu. "Economics Arena for Large Language Models." arXiv:2401.01735 (2024).