0

Circuit Supernodes

Fresh

Evidence/holdout causal-circuit, rediscovery, and restoration tasksets

Type
RL Env
Publisher
Wazupsteve
Runtime
multi-turn
License
unknown
Size
v0.1.1
Published
Oct 2026
Updated
Oct 2026

Cite

Notes

Only stored in your browser.

Bio-Circuit: causal circuit environments

Three text-only Verifiers v1 tasksets for interpreting Qwen3-4B through causal feature interventions. This development release provides public evaluation cohorts and source code. Scored rollouts require an external CUDA GPU backend. It is not a CPU-only environment and is not eligible for GPU-excluding runs.

Taskset IDTask
circuit-supernodesInfer a concept from evidence and identify causally necessary features
supernode-rediscoveryRediscover a circuit for a named semantic goal
suppressed-featuresRestore a subject model's output after feature corruption

All three tasksets and the circuit-null harness ship in this one Python wheel. The wheel includes backend source under modal_app; GPU dependencies are installed separately on the GPU host.

Public data

amitprakash2005/bio-circuit contains 16 evaluation tasks per taskset, full attribution graphs, provenance manifests, SHA-256 hashes, schemas, and lightweight browsable indexes. The data is CC-BY-4.0 and downloads without an HF token. The loader pins an immutable dataset commit; a supplied local dataset_file takes precedence.

Compressed data totals about 2.6 GiB, downloaded one taskset at a time. Parsed graphs need substantially more RAM. No training data or validated SFT trajectories are included. These evaluation cohorts are not a disjoint train set.

Install the environment client (no CUDA required)

Use Python 3.13 for the evaluation client:

uv venv --python 3.13
source .venv/bin/activate
uv pip install prime
prime env install wazupsteve/circuit-supernodes

The rediscovery/restoration modules and harness are included in that install. This package uses the Verifiers v1 Taskset interface with verifiers==0.2.1.

Start your own subject backend

The backend needs an 80 GB-class NVIDIA GPU, suitable CUDA drivers, and disk for the model and transcoders (allow at least 80 GB of cache/disk space). It loads public Qwen/Qwen3-4B and mwhanna/qwen3-4b-transcoders weights. Self-hosting requires no private service or author's Modal credentials.

Download and extract this environment's public source archive from Prime Hub. On the GPU host, from the extracted directory, install the backend in a separate Python 3.11 virtualenv. Do not install the evaluation client there; the backend and Verifiers use different numerical stacks.

uv venv --python 3.11 .venv-subject
uv pip install --python .venv-subject/bin/python \
  'torch==2.8.0' --index-url https://download.pytorch.org/whl/cu126
uv pip install --python .venv-subject/bin/python \
  'circuit-tracer==0.5.0' 'nnsight==0.6.1' 'transformer-lens==2.16.1' \
  'cloudpickle==3.0.0' 'huggingface-hub[hf_transfer]<1' \
  fastapi uvicorn typer pydantic tenacity httpx
.venv-subject/bin/python -m modal_app.server --port 8000

The server listens on loopback and serializes GPU operations. Reach it from the evaluation machine through an SSH tunnel:

ssh -N -L 8000:127.0.0.1:8000 user@gpu-host
# In another terminal with the evaluation client virtualenv activated:
export SUBJECT_URL=http://127.0.0.1:8000
curl --fail "$SUBJECT_URL/health"

The endpoint has no built-in authentication; keep it on loopback or a trusted private network. In a remote sandbox, localhost refers to that sandbox: configure a backend address reachable from both the tool runtime and reward process. The example below uses local subprocess runtimes.

Without SUBJECT_URL, the client uses the Modal app named circuit-supernodes. Deploy the included modal_app/attribution.py in your own Modal workspace and supply your own Modal credentials to use that option. This release does not provide a hosted GPU service or a published container image.

Run one evaluation

Save this as eval.toml and provide policy inference credentials through the provider configuration supported by Verifiers. Inference and GPU hosting incur costs.

model = "Qwen/Qwen3.5-4B"
num_tasks = 1
num_rollouts = 1
max_turns = 8

[sampling]
max_tokens = 16384
enable_thinking = false

[taskset]
id = "circuit-supernodes"

[harness]
id = "circuit-null"
eval @ eval.toml --dry-run --no-push
eval @ eval.toml --no-push

Change taskset.id to supernode-rediscovery or suppressed-features for the other tasks. Defaults select the corresponding public evaluation file. --dry-run checks configuration only; a real rollout establishes end-to-end backend operation. --no-push prevents result uploads, not provider calls.

Keep heldout/control cohorts and corruption specifications inaccessible to the policy. A public operator dataset is not a policy-visible tool response. These tasks have no gold-circuit labels; successful SFT trajectories must be generated and validated separately with live interventions and reward checks.

Development

Source: WazupSteve/bio-circuit. From environments/circuit_supernodes in the repository:

uv sync --frozen
uv run pytest
uv run ruff check .
uv build

The repository's configs/rl/ files reference training data not included in this release. Hosted training is not validated by local dry runs. Provision the backend and supply separately generated training data before attempting RL.