runebench
RuneBench on the Prime Environments Hub (verifiers v1): 40 RuneScape gameplay
tasks ported from MaxBittker/runebench
(MIT). The agent plays an emulated RuneScape private server (rs-sdk / LostCity
engine) running at 8x speed inside the sandbox, writing and executing
TypeScript bot scripts through the upstream rs-agent MCP server
(execute_code on bot/sdk globals), with the game wiki bundled in the
image for strategy reference.
- 32 skill-XP tasks — 16 skills x {15, 30} min. Reward: peak real-game XP/min over a single 15-second sampling window (raw XP / 8 game speed / 25 server XP rate; windows < 12 s ignored; tracker cadence 15 s).
- 8 gold tasks — 4 starting conditions (vanilla, smith-alch, fish, fletch-alch) x {15, 30} min. Reward: peak total coins across inventory + bank (tracker cadence 5 s; max of any sample's gold and the final save read).
Native verifiers v1 environment — no Harbor at runtime. The prime-agent harness (autonomous ACP) is the agent seat and consumes the rs-agent game tools as an MCP server.
Sandbox images
Thin layers over the upstream pre-built image
ghcr.io/maxbittker/rs-agent-benchmark:v71 (engine + gateway + bot client +
wiki + MCP server + skill tracker, with upstream's anti-tamper model intact:
root-owned engine/saves/tracker output, agent user unprivileged):
prime/primeintellect/runebench-skill:0.1.0—agent.savstart, 15 s tracker cadence (all 32 skill tasks)prime/primeintellect/runebench-gold-{vanilla,smith-alch,fish,fletch-alch}:0.1.0— per-condition starting save, 5 s tracker cadence (the 8 gold tasks)
Rebuild with prime images push runebench-<kind>:<tag> --context <dir>
(the image Dockerfiles live with this port; the base image is pinned, never
rebuilt).
Upstream fidelity and deviations
Prompts are the upstream generate-tasks.ts templates byte-for-byte (verified
against a fresh bun generate-tasks.ts run in tests). Scoring is a Python port
of shared/check_skill_xp.ts / shared/check_gold.ts with identical math and
constants. Deviations, all documented here:
- Scoring is native Python reward code, not the upstream TypeScript
verifier scripts run in-sandbox. Same constants and window semantics
(
MIN_PEAK_WINDOW_MS = 12000, /8 /25 normalization, rounded peak; peak-gold = max(final save, any sample)). Final xp/level metadata comes from the last tracker sample instead of a live gateway connection (upstream reads live state for metadata only; the reward math is tracker-derived in both). - 40 of 41 upstream tasks: the 5-minute woodcutting smoke task is not ported (upstream generates it as a harness sanity check, not a benchmark row).
- 5 images instead of 40: upstream generates one thin
FROMlayer per task (baking save +SAMPLE_INTERVAL_MS+BENCHMARK_DURATION_SECS); this port collapses to one image per (starting save, tracker cadence).BENCHMARK_DURATION_SECS(display-only in the agent-facingcheck_xp_rate.tsprogress line) moves to the task's runtime env. - MCP wiring: the upstream stdio MCP server (
bun run /app/mcp/server.ts) is proxied 1:1 to the harness through a colocated vf Toolset adapted from verifiers'HarborMCPToolset— the tool catalog, schemas, and the SDK API / MARKET resources are forwarded dynamically from the live upstream server. Harbor itself does not run. - Game-stack lifecycle: Prime VMs boot the image as a microVM rootfs under
/sbin/init— the Docker ENTRYPOINT does NOT run at boot (verified on the platform). The tasksetuptherefore starts/entrypoint.shitself when the stack is not already running (engine boot ~90 s on 2 vCPU), applies the upstream anti-cheat (save-generatorremoval), waits for the gateway, and runs the upstreamensure-services.sh(root-owned tracker up). The running-stack probe uses bracketed pgrep patterns (entrypoin[t].sh) — a barepgrep -f entrypoint.shinsidesh -cmatches the probe's own command line and reported a phantom running stack in the first hosted smoke. On stack failure, setup raises with the sandbox's process/port/log state embedded in the error.finalizeports the upstreamVERIFIER_CLEANUP(stop ffmpeg, kill orphaned agent bun scripts) before scoring. The setup timeout is 600 s (upstream Harbor allowedbuild_timeout = 1200): the stack boot and the harness's own in-sandbox install (node + prime-agent) share one setup deadline. - Harbor task metadata (author, difficulty, tags) is not carried; the upstream per-model agent adapters (opencode etc.) are replaced by the prime-agent harness seat.
Run
prime eval run primeintellect/runebench --hosted \
-m internal/glm-5.3-fast -n 40 -r 1 \
--env-args '{"agent": {"harness": {"id": "prime-agent", "autonomous": true}}}'
Attribution
Upstream: MaxBittker/runebench (MIT) and MaxBittker/rs-sdk (MIT), built on
LostCityRS/Server. The rs-agent MCP proxy adapts PrimeIntellect-ai/verifiers'
HarborMCPToolset (Apache-2.0).