0

Gtbench Eval

Fresh

GTBench economic game subset (auction, bargaining, Kuhn poker, liar's dice, iterated PD) as a verifiers v1 eval environment

Type
RL Env
License
mit
Size
v0.1.1
Published
Sep 2026
Updated
Sep 2026

Cite

Notes

Only stored in your browser.

gtbench-eval

GTBench ("GTBench: Uncovering the Strategic Reasoning Limitations of LLMs via Game-Theoretic Evaluations", NeurIPS 2024, https://github.com/jinhaoduan/GTBench) — economic game subset — as a verifiers v1 Environments-Hub eval environment. Vendored from the benchmark (MIT) with its engine re-implemented on open-spiel==1.4 (the benchmark's own pin) as a single-seat direct-model environment.

What the model under evaluation plays

One task = one OpenSpiel match driven with GTBench's own language protocol (system prompt, per-turn observation + step prompt, regex action parsing, last match wins). The model plays ONE seat; the other seat is a scripted opponent (no model calls). One verifiers interaction per episode carries the whole match transcript.

gameopponentwhat it measures
first_sealed_auctionequilibrium bidder b(v)=v/2bid shading vs. the FPSBA equilibrium
negotiationgreedy half-pie bargainermulti-issue bargaining: deals, own-value share
kuhn_pokerKuhn equilibrium mixture (α=1/3)imperfect-information betting
liars_diceminimum-raise bidderbluffing / probability under hidden dice
python_iterated_prisoners_dilemmatit-for-tatrepeated-game cooperation vs exploitation

Default suite: 5 games x 2 model seats x 3 seeds = 30 tasks (first_sealed_auction, negotiation, kuhn_poker, liars_dice, python_iterated_prisoners_dilemma). GTBench has no public-goods game; the non-economic board games (tictactoe, connect4, breakthrough, nim, pig) are out of scope per the port brief.

Scoring

Per-game components in [0,1] (composite = unweighted mean, recorded as reward gtbench weight 1.0 on the model seat trace; components at weight 0; all raw quantities + the full step log in trace.info["gtbench"]):

  • auction: valid (normal completion), won, surplus = (v − price)/v
  • negotiation: valid, deal, util_share = own deal utility / own pie
  • kuhn_poker: valid, win (1/0.5/0), payoff_norm = (r+2)/4
  • liars_dice: valid, win, payoff_norm = (r+1)/2
  • PD: valid, payoff_share = r/(r+o), abs_norm = r/(5·rounds)

An invalid/unparseable action ends the match Abnormal (GTBench semantics) and scores valid = 0.

Faithfulness notes (upstream quirks kept on purpose)

Ported verbatim from gamingbench/games/* and gamingbench/prompts/*:

  • kuhn_poker renders past-move roles with GTBench's modulo formula (attribution is wrong for seat 1 — upstream behavior);
  • the PD past-round rendering mirrors GTBench's (opponent history is rendered for both players);
  • negotiation accepts only the strict [a, b, c] bracket spacing of upstream's inner regex ([1,2,3] fails the move — benchmark-faithful);
  • liars_dice parses both dice from the debug state string (upstream does too) but the prompt only shows the model its own die;
  • negotiation runs with upstream's default (fixed) valuation vectors.

Deviations (documented): chance nodes sample from a dedicated np.random.default_rng(seed) (upstream uses the process-global numpy RNG); GTBench calls the model stateless per move, this port keeps one growing seat transcript per match (the per-turn prompts already render the benchmark's visible history); the LLM-vs-LLM win-rate/Elo scoring of the paper is replaced by the single-seat scripted-opponent components above.

Running

Local smoke (subprocess runtime, no Prime tunnel):

.venv/bin/eval gtbench_eval -m internal/glm-5.3-fast -n 2 -r 1 \
    --max-concurrent 1 --max-tokens 131072 --no-serve --no-rich \
    --env.agent.runtime.type subprocess \
    --env-args '{"taskset": {"games": ["first_sealed_auction"], "seeds": [8], "seats": [0]}}'

Hosted canary / probes:

prime eval run primeintellect/gtbench-eval --hosted \
    -m internal/glm-5.3-fast -n 1 -r 1 --max-concurrent 1 \
    --max-tokens 131072 --timeout-minutes 60 --eval-name gtbench-canary --plain

PrimeAgentHarness mode: pin the seat harness via --env-args '{"agent": {"harness": {"id": "prime-agent"}}}' (verified pattern from agent-bazaar-eval).

Model-free tests

.venv/bin/python -m pytest tests/ -q

Stub seats drive every game to terminal without model calls (canned legal actions), covering: match wiring, forced auto-moves, invalid-move Abnormal ends, the empty-reply guard, metric bounds, and the deterministic negotiation deal vs. the greedy opponent.

Traces + audit

scripts/audit_eval.py <eval-id> [--json out.json] folds traces into episodes per task and aggregates per game (means of composite + components), reporting harness id, stop conditions, and usage.