0

Emulatorbench V1

Fresh

Emulator Bench deterministic emulator implementation tasks for Verifiers v1.

Type
RL Env
Runtime
agent
License
unknown
Size
v0.1.12
Published
Sep 2026
Updated
Sep 2026

Cite

Notes

Only stored in your browser.

emulatorbench

Emulator Bench asks agents to implement deterministic Rust emulators for 16 classic computer and console platforms. This package contains the public native Verifiers v1 taskset, source manifests, starter workspace, feedback verifier, and scoring code. It is harness-agnostic and contains no private harness, credentials, private artifact payloads, or private download locations.

Canonical interface

  • Distribution and taskset ID: emulatorbench
  • Python module: emulatorbench
  • Intended Environments Hub identity: primeintellect/emulatorbench
  • Task data/config classes: EmulatorBenchData, EmulatorBenchTaskConfig, and EmulatorBenchConfig
  • Task class: EmulatorBenchTask
  • Taskset class: EmulatorBenchTaskset
  • Task/platform IDs: emulatorbench-<platform>
  • Sandbox label metadata: exactly emulatorbench
  • Standard Prime Agent evaluation: PRIME_EVAL.md, using PLATFORM=<slug> ./scripts/run_prime_eval.sh

The runtime image is configurable with --taskset.image or EMULATORBENCH_IMAGE. The default is the public digest-pinned rust:1.85-bookworm@sha256:e51d0265072d2d9d5d320f6a44dde6b9ef13653b035098febd68cce8fa7c0bc4 image. External-controller grading rejects mutable image references. Runtime and harness selection remain caller-owned.

Platforms

Platform slugTask IDSystemLevelDeclared units
chip8emulatorbench-chip8CHIP-818
i8080_space_invadersemulatorbench-i8080-space-invadersIntel 8080 / Space Invaders25
gameboy_dmgemulatorbench-gameboy-dmgNintendo Game Boy (DMG)32,580,686
nesemulatorbench-nesNintendo Entertainment System4293
smsemulatorbench-smsSega Master System51,604,017
gameboy_cgbemulatorbench-gameboy-cgbGame Boy Color62,584,110
gbaemulatorbench-gbaGame Boy Advance757,571
genesisemulatorbench-genesisSega Genesis / Mega Drive82,604,060
snesemulatorbench-snesSuper Nintendo95,376,331
ps1emulatorbench-ps1Sony PlayStation10124
c64emulatorbench-c64Commodore 64111,544
n64emulatorbench-n64Nintendo 6412698
pspemulatorbench-pspPlayStation Portable13433
ps2emulatorbench-ps2PlayStation 21478
zx_spectrumemulatorbench-zx-spectrumZX Spectrum 48K151,605,358
game_gearemulatorbench-game-gearSega Game Gear161,604,027

Source and score integrity

Git sources are pinned to full commits and fetched with an exact-commit shallow fetch. Downloadable HTTP artifacts are SHA-256 pinned. Documentation references are not treated as executable artifacts. Package load validates IDs, source coverage, private descriptor boundaries, unique case metadata, and declared unit counts.

Every scored declaration has an explicit expected-unit denominator. Missing observations, inputs, or oracles never shrink that denominator. Runner-v1 output is retained only as an untrusted development diagnostic and always receives zero trusted reward. Positive reward requires a deeply validated runner-v2 response, a signed controller-owned input/oracle corpus, and a score record bound to that corpus digest and version.

The public taskset does not accept private artifact paths. Private payloads and oracles remain external-controller extensions. Public manifests contain non-payload descriptors only. Candidate and grader workspaces, submissions, results, lifecycle events, patches, and logs are retained under restricted host directories with credential exclusion and text redaction before teardown.

Feedback boundary

Same-UID public feedback is either the explicit full verifier or disabled. Aggregate and hidden modes are rejected because a candidate workspace cannot hide verifier inputs or enforce budgets.

With anti_cheat_validation, the candidate workspace receives no verifier, source manifest, expected output, private payload, or signing material. When public feedback is enabled, a host watcher transfers deterministic candidate archives to a distinct digest-pinned no-egress grader, then publishes bounded controller-signed count, diagnostic, or full feedback. Records bind task, trace, attempt, mode, raw archive, sanitized submission, retained grader result, and oracle corpus. The candidate helper verifies the signature and exact raw archive before displaying feedback. Disabling public feedback keeps the same signed final-grade path without staging a helper or starting a watcher.

Every grader gets a distinct provider identity, a bounded start, staging, scoring, capture, and confirmed teardown lifecycle, a provider hard-lifetime backstop, and restricted pre-teardown artifacts that exclude trusted /tmp verifier and oracle files. Any identity, egress, signing, artifact, lifecycle, result-digest, corpus, or confirmed-teardown failure zeros trusted reward.

emulator-runner-v2 accepts input-only ordered requests and returns bounded raw observations without pass or score claims. The controller keeps expected states and comparisons separate. The package currently includes strict schemas, signed-corpus verification, fixed plans for all 13 source-separable declarations, and fail-closed classifications for all 65 declarations. No declaration is marked ported until its complete signed corpus has been built and installed.

Safe validation

Import and construction do not launch a runtime:

python - <<'PY'
from verifiers.v1.loaders import taskset_class, taskset_config_type
print(taskset_class("emulatorbench"))
print(taskset_config_type("emulatorbench"))
PY

A real evaluation requires an isolated container runtime. Do not launch one without an explicit run plan.

Provenance reconciliation

PROVENANCE_RECONCILIATION.md records the reviewed implementation origin, intentional cleanup differences, and the handoff point for the CPU-run source comparison. Behavioral differences not listed there require reconciliation before approval.

Changelog

  • 2026-07-16: Added the canonical public Emulator Bench taskset from the reviewed PR #662 implementation, removed legacy harness/runtime compatibility paths, pinned public sources, corrected scoring denominators/composition, and added redacted terminal artifact capture.
  • 2026-07-17: Restored the external signed-controller feedback loop, added fresh confirmed-teardown graders and complete restricted artifact retention, bound scores to signed runner-v2 corpora, and made every unported runner-v1 declaration fail closed.
  • 2026-07-24: Ported chip8:chip8_timendus_suite as the first signed_corpus_installed runner-v2 plan, with immutable Timendus fixtures, a digest-pinned packaged verification key, bounded signed feedback, and lifecycle-gated trusted scoring.

Deviations

  • Offline-build disclosure (added 2026-09-23, v0.1.10): the task instruction now discloses that the grading build runs fully offline (cargo build --locked, no crate fetches) and that external crates must be vendored or avoided. Rationale: the original instruction invited adding dependencies without disclosing the offline grading constraint, and the candidate sandbox's network access let in-session builds pass while only the no-egress grader failed — an invisible trap (observed: glm psp 0.0 in run 008; gpt-6-astra systematic 0.0s on chip8/sms/psp with in-session public scores up to 99.9989%). The scoring contract is unchanged; the fix makes the constraint agent-visible.