emulatorbench
Emulator Bench asks agents to implement deterministic Rust emulators for 16 classic computer and console platforms. This package contains the public native Verifiers v1 taskset, source manifests, starter workspace, feedback verifier, and scoring code. It is harness-agnostic and contains no private harness, credentials, private artifact payloads, or private download locations.
Canonical interface
- Distribution and taskset ID:
emulatorbench - Python module:
emulatorbench - Intended Environments Hub identity:
primeintellect/emulatorbench - Task data/config classes:
EmulatorBenchData,EmulatorBenchTaskConfig, andEmulatorBenchConfig - Task class:
EmulatorBenchTask - Taskset class:
EmulatorBenchTaskset - Task/platform IDs:
emulatorbench-<platform> - Sandbox label metadata: exactly
emulatorbench - Standard Prime Agent evaluation:
PRIME_EVAL.md, usingPLATFORM=<slug> ./scripts/run_prime_eval.sh
The runtime image is configurable with --taskset.image or
EMULATORBENCH_IMAGE. The default is the public digest-pinned
rust:1.85-bookworm@sha256:e51d0265072d2d9d5d320f6a44dde6b9ef13653b035098febd68cce8fa7c0bc4 image. External-controller grading rejects mutable image
references. Runtime and harness selection remain caller-owned.
Platforms
| Platform slug | Task ID | System | Level | Declared units |
|---|---|---|---|---|
chip8 | emulatorbench-chip8 | CHIP-8 | 1 | 8 |
i8080_space_invaders | emulatorbench-i8080-space-invaders | Intel 8080 / Space Invaders | 2 | 5 |
gameboy_dmg | emulatorbench-gameboy-dmg | Nintendo Game Boy (DMG) | 3 | 2,580,686 |
nes | emulatorbench-nes | Nintendo Entertainment System | 4 | 293 |
sms | emulatorbench-sms | Sega Master System | 5 | 1,604,017 |
gameboy_cgb | emulatorbench-gameboy-cgb | Game Boy Color | 6 | 2,584,110 |
gba | emulatorbench-gba | Game Boy Advance | 7 | 57,571 |
genesis | emulatorbench-genesis | Sega Genesis / Mega Drive | 8 | 2,604,060 |
snes | emulatorbench-snes | Super Nintendo | 9 | 5,376,331 |
ps1 | emulatorbench-ps1 | Sony PlayStation | 10 | 124 |
c64 | emulatorbench-c64 | Commodore 64 | 11 | 1,544 |
n64 | emulatorbench-n64 | Nintendo 64 | 12 | 698 |
psp | emulatorbench-psp | PlayStation Portable | 13 | 433 |
ps2 | emulatorbench-ps2 | PlayStation 2 | 14 | 78 |
zx_spectrum | emulatorbench-zx-spectrum | ZX Spectrum 48K | 15 | 1,605,358 |
game_gear | emulatorbench-game-gear | Sega Game Gear | 16 | 1,604,027 |
Source and score integrity
Git sources are pinned to full commits and fetched with an exact-commit shallow fetch. Downloadable HTTP artifacts are SHA-256 pinned. Documentation references are not treated as executable artifacts. Package load validates IDs, source coverage, private descriptor boundaries, unique case metadata, and declared unit counts.
Every scored declaration has an explicit expected-unit denominator. Missing observations, inputs, or oracles never shrink that denominator. Runner-v1 output is retained only as an untrusted development diagnostic and always receives zero trusted reward. Positive reward requires a deeply validated runner-v2 response, a signed controller-owned input/oracle corpus, and a score record bound to that corpus digest and version.
The public taskset does not accept private artifact paths. Private payloads and oracles remain external-controller extensions. Public manifests contain non-payload descriptors only. Candidate and grader workspaces, submissions, results, lifecycle events, patches, and logs are retained under restricted host directories with credential exclusion and text redaction before teardown.
Feedback boundary
Same-UID public feedback is either the explicit full verifier or disabled.
Aggregate and hidden modes are rejected because a candidate workspace cannot
hide verifier inputs or enforce budgets.
With anti_cheat_validation, the candidate workspace receives no verifier,
source manifest, expected output, private payload, or signing material. When
public feedback is enabled, a host watcher transfers deterministic candidate
archives to a distinct digest-pinned no-egress grader, then publishes bounded
controller-signed count, diagnostic, or full feedback. Records bind task, trace,
attempt, mode, raw archive, sanitized submission, retained grader result, and
oracle corpus. The candidate helper verifies the signature and exact raw archive
before displaying feedback. Disabling public feedback keeps the same signed
final-grade path without staging a helper or starting a watcher.
Every grader gets a distinct provider identity, a bounded start, staging,
scoring, capture, and confirmed teardown lifecycle, a provider hard-lifetime
backstop, and restricted pre-teardown artifacts that exclude trusted /tmp
verifier and oracle files. Any identity, egress, signing, artifact, lifecycle,
result-digest, corpus, or confirmed-teardown failure zeros trusted reward.
emulator-runner-v2 accepts input-only ordered requests and returns bounded raw
observations without pass or score claims. The controller keeps expected states
and comparisons separate. The package currently includes strict schemas,
signed-corpus verification, fixed plans for all 13 source-separable declarations,
and fail-closed classifications for all 65 declarations. No declaration is
marked ported until its complete signed corpus has been built and installed.
Safe validation
Import and construction do not launch a runtime:
python - <<'PY'
from verifiers.v1.loaders import taskset_class, taskset_config_type
print(taskset_class("emulatorbench"))
print(taskset_config_type("emulatorbench"))
PY
A real evaluation requires an isolated container runtime. Do not launch one without an explicit run plan.
Provenance reconciliation
PROVENANCE_RECONCILIATION.md records the reviewed implementation origin,
intentional cleanup differences, and the handoff point for the CPU-run source
comparison. Behavioral differences not listed there require reconciliation
before approval.
Changelog
- 2026-07-16: Added the canonical public Emulator Bench taskset from the reviewed PR #662 implementation, removed legacy harness/runtime compatibility paths, pinned public sources, corrected scoring denominators/composition, and added redacted terminal artifact capture.
- 2026-07-17: Restored the external signed-controller feedback loop, added fresh confirmed-teardown graders and complete restricted artifact retention, bound scores to signed runner-v2 corpora, and made every unported runner-v1 declaration fail closed.
- 2026-07-24: Ported
chip8:chip8_timendus_suiteas the firstsigned_corpus_installedrunner-v2 plan, with immutable Timendus fixtures, a digest-pinned packaged verification key, bounded signed feedback, and lifecycle-gated trusted scoring.
Deviations
- Offline-build disclosure (added 2026-09-23, v0.1.10): the task instruction now
discloses that the grading build runs fully offline (
cargo build --locked, no crate fetches) and that external crates must be vendored or avoided. Rationale: the original instruction invited adding dependencies without disclosing the offline grading constraint, and the candidate sandbox's network access let in-session builds pass while only the no-egress grader failed — an invisible trap (observed: glm psp 0.0 in run 008; gpt-6-astra systematic 0.0s on chip8/sms/psp with in-session public scores up to 99.9989%). The scoring contract is unchanged; the fix makes the constraint agent-visible.