0

Gamedevbench

Fresh

GameDevBench: 333 Godot 4.4.1 game-development tasks for coding agents (eval-only; ground truths are public - do not train on these tasks)

Type
RL Env
Runtime
multi-turn
License
unknown
Size
v0.1.5
Published
Sep 2026
Updated
Sep 2026

Cite

Notes

Only stored in your browser.

GameDevBench — Prime Environments Hub package (eval-only)

Port of https://github.com/waynchi/gamedevbench (ICML 2026): 333 Godot 4.4.1 game-development tasks, scored by the benchmark's official headless Godot validation (per-task scripts/test.gd + scenes/test.tscn printing VALIDATION_PASSED / VALIDATION_FAILED).

EVAL-ONLY: ground truths and tutorial sources are public. Do not train on these tasks.

Layout

  • gamedevbench/taskset.py — verifiers v1 Taskset + Task (setup + @reward scoring that mirrors the benchmark's strict-confinement validation flow)
  • gamedevbench/tasks_meta.json — per-task instruction / display flag
  • Dockerfile — task image: Godot 4.4.1 (exact), xvfb/xauth, the 333 official task zips at /opt/gamedevbench/tasks (pinned to gamedevbench@3a0dd3c)

Deviations from the official benchmark (documented)

  1. Prompt, text-only models: the official text-mode prompt line "You are a visual agent and can use images and videos to help you understand the state of the game." is replaced by an explicit text-only instruction ("Do not use attach_image, generate screenshots, or create movie files. Verify ... through headless Godot command output, print statements, and text logs."). Reason: with glm-5.3-fast (no vision), the official line drives agents to attach screenshots; requests carrying image attachments upstream-500 and kill the episode after one retry (verified on two hosted canaries, 2026-09-21). All other prompt bytes are identical to gamedevbench/src/utils/prompts.py::create_task_prompt with use_runtime_video=False, use_mcp=False.
  2. Harness binding: the env config binds the agent seat to vf.AgentConfig(harness={"id": "prime-agent", "autonomous": True}). glm-5.3-fast's first turn is often reasoning-only with no tool calls; non-autonomous rollouts end after one call (verified on canary 1). Autonomous continuation keeps rollouts alive until task completion or the 1800s agent timeout.

Flow per task

  1. setup(): unzip the task into /workspace under the official sandbox filter (no test*, task_config.json, *.log, *.md, hidden files/dirs), write the minimal task_config.json, headless --import.
  2. Prime Agent (ACP) solves the task in /workspace with the official text-mode prompt.
  3. godot_validation reward: copy the workspace, restore test.gd/test.tscn from the pristine zip, headless --import, run scenes/test.tscn (xvfb-run for the 10 display-required tasks), parse VALIDATION_PASSED.

Run

prime images push gamedevbench:0.1.0 --path . --dockerfile Dockerfile --plain
prime env push --path . --name gamedevbench --visibility PRIVATE --runtime v1 --plain
prime eval run primeintellect/gamedevbench --hosted \
    -m internal/glm-5.3-fast -n 333 -r 1 --max-concurrent 32 \
    --max-tokens 131072 --timeout-minutes 300 \
    --eval-name gamedevbench-glm53fast-001 --plain