0

LLM Economist Eval

Fresh

LLM Economist Stackelberg tax-policy benchmark as a verifiers v1 eval environment

Type
RL Env
License
mit
Size
v0.1.4
Published
Sep 2026
Updated
Sep 2026

Cite

Notes

Only stored in your browser.

llm-economist-eval

LLM Economist (Stackelberg tax-policy world, https://github.com/sethkarten/LLM-Economist, arXiv:2507.15815) as a verifiers v1 Environments-Hub eval environment. Vendored from the MIT-licensed upstream repo.

What the model under evaluation plays

Each task = one full tax-policy episode for one (scenario cell, seed); 6 cells x 3 seeds = 18 tasks. One verifiers interaction (seat) per LLM sim agent; each benchmark decision prompt is one seat turn.

  • stackelberg_llm: the TaxPlanner seat (bracket-delta tax schedules every tax year from worker (z, utility) histograms) + all 5 LLM worker seats (weekly labor hours in [0,100]).
  • planner_fixedpop: only the planner seat; procedural FixedWorkers.
  • workers_saez / workers_usfed: only worker seats, governed by the analytic Saez schedule (three brackets, elasticity 3.0) or the statutory 2024 US federal schedule (seven brackets).
  • bounded_personas: planner + persona worker seats; workers answer a per-step satisfaction reflection (utility scaled by 1.0 YES / 0.5 NO).
  • democratic_platforms: worker seats only; every tax year workers campaign with a platform (proposed deltas), vote, and the elected leader sets the tax deltas.

Rewards recorded per seat trace: econ (composite, weight 1) + econ_welfare / econ_productivity / econ_equity (weight 0), with raw quantities and per-step series in trace.info["llm_econ"].

Metrics

All components in [0,1], higher = better, evaluated over the terminal tax year (see metrics.py docstring for the definitions and caveats):

  • welfare: mean over agents of clip01(u / max(z, 1)) — the benchmark's equity-weighted SWF per capita, per-agent clipped at 1 (unclipped SWF in trace.info raw).
  • productivity: mean labor / 100.
  • equity: 1 - Gini(consumption) at episode end (consumption = post-tax income + rebate).

Composite = unweighted mean of the three.

Scenario cells (experiment fidelity)

Cells mirror experiments/run_experiments.py base configs (rational scenario, us_income GB2 skills, history-len 50, io prompts, temperature 0.7). Scaling deviation, documented per the paper: the paper's central config is N=100 / T=3000 / K=128 tax years; hosted full-effort episodes must stay eval-sized, so cells default to N=5 / T=50 / K=25 (two tax years = one in-context planner revision per episode). Vendored patches: wandb/torch stripped; provider model clients lazily imported; the fixed-persona path returns a proper dict (upstream returned a bare list, which crashed GEN_ROLE_MESSAGES.update) and skips the census CSV that the upstream repo does not ship; --timeout (JSON-retry budget per decision, vendored default 10) is actually forwarded to agents (upstream parses it but drops it) and pinned to 3 in all cells.

Reproducibility

Seeds are applied exactly as llm_economist.main does (numpy + python globals before construction). The benchmark draws from process-global RNGs during stepping and shares the module-global GEN_ROLE_MESSAGES persona map, so run episodes one-per-process for exact reproducibility: --serve.max-concurrent 1, or --max-concurrent 1 in-process. Throughput then scales with pool workers (processes).

Running

Local:

.venv/bin/eval llm_economist_eval --env-dir-path . \
    -m internal/glm-5.3-fast --env.agent.runtime.type subprocess \
    -n 1 -r 1 --max-concurrent 1 --max-tokens 131072 --no-serve --no-rich

Hosted (canary first; probe overrides via --env-args):

prime eval run primeintellect/llm-economist-eval --hosted \
    -m internal/glm-5.3-fast -n 1 -r 1 --max-concurrent 1 \
    --max-tokens 131072 --timeout-minutes 60 \
    --eval-name econ-canary --plain \
    --env-args '{"taskset": {"cells": ["stackelberg_llm"], "seeds": [8],
                              "timesteps": 2, "two_timescale": 1}}'

PrimeAgentHarness mode: add "agent": {"harness": {"id": "prime-agent"}} to --env-args (and pin "max_concurrent_agents": 4, "runtime": {"cpu": 8, "memory": 8} for seat-heavy episodes).

Vendoring notes (llm_economist)

  • llm_economist/ is vendored top-level (absolute llm_economist.* imports).
  • Patched: main.py strips wandb + the torch bootstrap (kept only for create_argument_parser()); models/__init__.py ships only BaseLLMModel; agents/llm_agent.py + agents/worker.py import provider clients lazily; agents/worker.py::distribute_personas fixed-persona path returns {persona_i: description} without the missing census CSV.
  • The eval env drives the episode loop itself (env.py::_drive mirrors run_simulation step by step); main.run_simulation is never called.