gtbench-eval
GTBench ("GTBench: Uncovering the Strategic Reasoning Limitations of LLMs
via Game-Theoretic Evaluations", NeurIPS 2024,
https://github.com/jinhaoduan/GTBench) — economic game subset — as a
verifiers v1 Environments-Hub eval environment. Vendored from the
benchmark (MIT) with its engine re-implemented on open-spiel==1.4
(the benchmark's own pin) as a single-seat direct-model environment.
What the model under evaluation plays
One task = one OpenSpiel match driven with GTBench's own language protocol (system prompt, per-turn observation + step prompt, regex action parsing, last match wins). The model plays ONE seat; the other seat is a scripted opponent (no model calls). One verifiers interaction per episode carries the whole match transcript.
| game | opponent | what it measures |
|---|---|---|
first_sealed_auction | equilibrium bidder b(v)=v/2 | bid shading vs. the FPSBA equilibrium |
negotiation | greedy half-pie bargainer | multi-issue bargaining: deals, own-value share |
kuhn_poker | Kuhn equilibrium mixture (α=1/3) | imperfect-information betting |
liars_dice | minimum-raise bidder | bluffing / probability under hidden dice |
python_iterated_prisoners_dilemma | tit-for-tat | repeated-game cooperation vs exploitation |
Default suite: 5 games x 2 model seats x 3 seeds = 30 tasks
(first_sealed_auction, negotiation, kuhn_poker, liars_dice,
python_iterated_prisoners_dilemma). GTBench has no public-goods
game; the non-economic board games (tictactoe, connect4, breakthrough,
nim, pig) are out of scope per the port brief.
Scoring
Per-game components in [0,1] (composite = unweighted mean, recorded as
reward gtbench weight 1.0 on the model seat trace; components at
weight 0; all raw quantities + the full step log in
trace.info["gtbench"]):
- auction:
valid(normal completion),won,surplus= (v − price)/v - negotiation:
valid,deal,util_share= own deal utility / own pie - kuhn_poker:
valid,win(1/0.5/0),payoff_norm= (r+2)/4 - liars_dice:
valid,win,payoff_norm= (r+1)/2 - PD:
valid,payoff_share= r/(r+o),abs_norm= r/(5·rounds)
An invalid/unparseable action ends the match Abnormal (GTBench
semantics) and scores valid = 0.
Faithfulness notes (upstream quirks kept on purpose)
Ported verbatim from gamingbench/games/* and gamingbench/prompts/*:
- kuhn_poker renders past-move roles with GTBench's modulo formula (attribution is wrong for seat 1 — upstream behavior);
- the PD past-round rendering mirrors GTBench's (opponent history is rendered for both players);
- negotiation accepts only the strict
[a, b, c]bracket spacing of upstream's inner regex ([1,2,3]fails the move — benchmark-faithful); - liars_dice parses both dice from the debug state string (upstream does too) but the prompt only shows the model its own die;
- negotiation runs with upstream's default (fixed) valuation vectors.
Deviations (documented): chance nodes sample from a dedicated
np.random.default_rng(seed) (upstream uses the process-global numpy
RNG); GTBench calls the model stateless per move, this port keeps one
growing seat transcript per match (the per-turn prompts already render
the benchmark's visible history); the LLM-vs-LLM win-rate/Elo scoring of
the paper is replaced by the single-seat scripted-opponent components
above.
Running
Local smoke (subprocess runtime, no Prime tunnel):
.venv/bin/eval gtbench_eval -m internal/glm-5.3-fast -n 2 -r 1 \
--max-concurrent 1 --max-tokens 131072 --no-serve --no-rich \
--env.agent.runtime.type subprocess \
--env-args '{"taskset": {"games": ["first_sealed_auction"], "seeds": [8], "seats": [0]}}'
Hosted canary / probes:
prime eval run primeintellect/gtbench-eval --hosted \
-m internal/glm-5.3-fast -n 1 -r 1 --max-concurrent 1 \
--max-tokens 131072 --timeout-minutes 60 --eval-name gtbench-canary --plain
PrimeAgentHarness mode: pin the seat harness via --env-args '{"agent": {"harness": {"id": "prime-agent"}}}' (verified pattern from
agent-bazaar-eval).
Model-free tests
.venv/bin/python -m pytest tests/ -q
Stub seats drive every game to terminal without model calls (canned legal actions), covering: match wiring, forced auto-moves, invalid-move Abnormal ends, the empty-reply guard, metric bounds, and the deterministic negotiation deal vs. the greedy opponent.
Traces + audit
scripts/audit_eval.py <eval-id> [--json out.json] folds traces into
episodes per task and aggregates per game (means of composite +
components), reporting harness id, stop conditions, and usage.