llm-economist-eval
LLM Economist (Stackelberg tax-policy world, https://github.com/sethkarten/LLM-Economist, arXiv:2507.15815) as a verifiers v1 Environments-Hub eval environment. Vendored from the MIT-licensed upstream repo.
What the model under evaluation plays
Each task = one full tax-policy episode for one (scenario cell, seed); 6 cells x 3 seeds = 18 tasks. One verifiers interaction (seat) per LLM sim agent; each benchmark decision prompt is one seat turn.
- stackelberg_llm: the TaxPlanner seat (bracket-delta tax schedules every tax year from worker (z, utility) histograms) + all 5 LLM worker seats (weekly labor hours in [0,100]).
- planner_fixedpop: only the planner seat; procedural FixedWorkers.
- workers_saez / workers_usfed: only worker seats, governed by the analytic Saez schedule (three brackets, elasticity 3.0) or the statutory 2024 US federal schedule (seven brackets).
- bounded_personas: planner + persona worker seats; workers answer a per-step satisfaction reflection (utility scaled by 1.0 YES / 0.5 NO).
- democratic_platforms: worker seats only; every tax year workers campaign with a platform (proposed deltas), vote, and the elected leader sets the tax deltas.
Rewards recorded per seat trace: econ (composite, weight 1) +
econ_welfare / econ_productivity / econ_equity (weight 0), with raw
quantities and per-step series in trace.info["llm_econ"].
Metrics
All components in [0,1], higher = better, evaluated over the terminal tax
year (see metrics.py docstring for the definitions and caveats):
- welfare: mean over agents of clip01(u / max(z, 1)) — the benchmark's equity-weighted SWF per capita, per-agent clipped at 1 (unclipped SWF in trace.info raw).
- productivity: mean labor / 100.
- equity: 1 - Gini(consumption) at episode end (consumption = post-tax income + rebate).
Composite = unweighted mean of the three.
Scenario cells (experiment fidelity)
Cells mirror experiments/run_experiments.py base configs (rational
scenario, us_income GB2 skills, history-len 50, io prompts, temperature
0.7). Scaling deviation, documented per the paper: the paper's central
config is N=100 / T=3000 / K=128 tax years; hosted full-effort episodes must
stay eval-sized, so cells default to N=5 / T=50 / K=25 (two tax years = one
in-context planner revision per episode). Vendored patches: wandb/torch
stripped; provider model clients lazily imported; the fixed-persona path
returns a proper dict (upstream returned a bare list, which crashed
GEN_ROLE_MESSAGES.update) and skips the census CSV that the upstream repo
does not ship; --timeout (JSON-retry budget per decision, vendored default
10) is actually forwarded to agents (upstream parses it but drops it) and
pinned to 3 in all cells.
Reproducibility
Seeds are applied exactly as llm_economist.main does (numpy + python
globals before construction). The benchmark draws from process-global RNGs
during stepping and shares the module-global GEN_ROLE_MESSAGES persona map,
so run episodes one-per-process for exact reproducibility:
--serve.max-concurrent 1, or --max-concurrent 1 in-process. Throughput
then scales with pool workers (processes).
Running
Local:
.venv/bin/eval llm_economist_eval --env-dir-path . \
-m internal/glm-5.3-fast --env.agent.runtime.type subprocess \
-n 1 -r 1 --max-concurrent 1 --max-tokens 131072 --no-serve --no-rich
Hosted (canary first; probe overrides via --env-args):
prime eval run primeintellect/llm-economist-eval --hosted \
-m internal/glm-5.3-fast -n 1 -r 1 --max-concurrent 1 \
--max-tokens 131072 --timeout-minutes 60 \
--eval-name econ-canary --plain \
--env-args '{"taskset": {"cells": ["stackelberg_llm"], "seeds": [8],
"timesteps": 2, "two_timescale": 1}}'
PrimeAgentHarness mode: add "agent": {"harness": {"id": "prime-agent"}}
to --env-args (and pin "max_concurrent_agents": 4, "runtime": {"cpu": 8, "memory": 8} for seat-heavy episodes).
Vendoring notes (llm_economist)
llm_economist/is vendored top-level (absolutellm_economist.*imports).- Patched:
main.pystrips wandb + the torch bootstrap (kept only forcreate_argument_parser());models/__init__.pyships onlyBaseLLMModel;agents/llm_agent.py+agents/worker.pyimport provider clients lazily;agents/worker.py::distribute_personasfixed-persona path returns{persona_i: description}without the missing census CSV. - The eval env drives the episode loop itself (
env.py::_drivemirrorsrun_simulationstep by step);main.run_simulationis never called.