SMDD-Bench on Prime
Paper | SMDD-Bench website | Blog
SMDD-Bench tests whether language-model agents can carry out long-horizon small-molecule drug design. This package is the Prime-native release: it bundles the task data, original agent harness, scoring code, and Docker build sources needed to run SMDD-Bench with Prime's verifiers.v1 evaluation runner.
There are five task families:
| Type | Task | Submission |
|---|---|---|
| 1 | 2D Pharmacophore Identification | solution.py |
| 2 | Interaction Point Discovery | solution.csv |
| 3 | Scaffold Hopping | solution.smi |
| 4 | Lead Optimization | solution.smi |
| 5 | Fragment Assembly | solution.smi |
The full benchmark contains 502 tasks. Some validation builds contain only a subset; run smdd-bench list to see exactly what is installed in your release.
How the setup fits together
The deployment has two moving parts:
The evaluation host runs Prime, calls the language model, and launches the task and verifier containers.
The oracle service runs Boltz and ADMET on GPU. It can live on the same machine or on a separate GPU machine / Ray cluster.
Prime handles the short-lived task and verifier containers for you. The oracle is the only long-running service you start yourself.
The task and verifier are intentionally separated. Public task files are staged into the agent runtime; only declared submission artifacts are copied into the verifier runtime, where the private grader inputs are staged. The operator still has access to the release files, but the agent does not receive the grader, oracle token, or host model credentials.
Before you start
Use a Linux host with:
- Docker
- Python 3.12
uv
Every GPU host also needs the NVIDIA driver and NVIDIA Container Toolkit. Before building anything, confirm that nvidia-smi works and that Docker can access the GPU.
For Boltz, we recommend at least 48 GB of VRAM per GPU to reduce out-of-memory failures. Actual memory use depends on the protein and inference settings, so 48 GB is not a guarantee for every task. For larger workloads, an 8 x A100 80 GB machine is one option.
Multiple GPUs increase throughput by serving independent requests. Their memory is not automatically pooled into one larger device for a single prediction. Start with one trial, watch GPU and host memory, and only then increase concurrency.
You will also need enough disk space for Docker images and model checkpoints. The oracle requires outbound network access to its MSA service. The evaluation host needs network access to both your model provider and the oracle endpoint.
1. Install the Prime environment
Run the following in Bash on the Linux evaluation host.
The package is published as looni-lab/smdd-bench. Version 0.1.4.post0 is currently private, so log in with a Prime account that has access.
uv venv --python 3.12 "$HOME/.venvs/smdd-bench"
source "$HOME/.venvs/smdd-bench/bin/activate"
command -v prime >/dev/null || uv tool install prime
prime login
prime env install looni-lab/smdd-bench@0.1.4.post0
Prime may install the downloaded wheel into its own tool Python even when a virtual environment is active. Check the install output. To make sure the release is available inside the evaluation environment, install the downloaded wheel explicitly:
WHEEL="$HOME/.prime/wheel_cache/looni-lab/smdd-bench/0.1.4.post0/dist/smdd_bench-0.1.4.post0-py3-none-any.whl"
uv pip install --python "$VIRTUAL_ENV/bin/python" "$WHEEL"
If Prime printed a different wheel path, use that path instead.
Now check the installed task set and create a writable run directory:
smdd-bench list
smdd-bench setup "$HOME/smdd-bench-run"
cd "$HOME/smdd-bench-run"
No repository checkout is required; the package already includes its task data and runtime sources.
When migrating from an older SMDD distribution, use the new virtual environment above and generate a fresh run directory with smdd-bench setup; old TOML configs still reference the previous loader IDs.
Keep this Python environment activated while running SMDD. Avoid installing another SMDD distribution into the same environment, since different releases may contain incompatible versions of the shared harness module.
If smdd-bench: command not found appears, first check that the virtual environment is still active and that the explicit wheel installation succeeded.
Some Prime CLI versions print a generic from verifiers import load_environment example. That uses the older API. This release exports native classes, so use the smdd-bench eval commands in this guide instead.
For these VM-based evaluations, you do not need Prime Hub Secrets or Variables. Model and oracle credentials are configured in the VM shell below; Hub settings are for hosted services and do not configure this local runner.
What was installed
The Python package contains the canonical task definitions, harness, runtime sources, and task data:
smdd_bench/
|-- taskset.py # Canonical tasks and native Prime loading
|-- harness.py # Reference runs and model-backed agent runs
|-- worker.py, relay.py # Host agent loop and isolated task operations
|-- release.json # Task IDs, image tags and source hashes
|-- assets/smdd-runtime.zip # Docker and oracle build sources
+-- tasks/<task-id>/
|-- task.toml # Resources, images and separate verifier policy
|-- instruction.md # Public task instructions
|-- environment/ # Public files placed in the agent workspace
|-- tests/ # Grader and private scoring inputs
+-- solution/ # Reference solution, when available
smdd-bench setup creates a separate writable run directory:
smdd-bench-run/
|-- release.json
|-- lead.toml # One real lead-optimization trial
|-- lead-smoke.toml # Lead reference solution and verifier
|-- smoke.toml # One reference task from each family
+-- smdd-runtime/ # Extracted sources, including build_images.sh
2. Build the Docker images
The release ships Docker build sources rather than prebuilding everything on your machine. build_images.sh uses the exact image tags recorded in release.json.
| Image role | Build on | Used for |
|---|---|---|
| Agent science and evaluator images | Evaluation host | Parent images for task and verifier containers |
| Task science and task lead images | Evaluation host | Agent workspaces |
| Five verifier images | Evaluation host | Separate scoring environments for the five task families |
| Oracle model image | GPU build host | Boltz, ADMET, and Ray dependencies |
| HTTP oracle image | Every GPU VM | Long-running oracle service |
Choose the build that matches your deployment:
# Evaluation and oracle on this same GPU host:
bash smdd-runtime/build_images.sh all
# Alternatively, evaluation host only, or reusing an existing GPU oracle:
bash smdd-runtime/build_images.sh tasks
# Separate GPU host, after extracting the same package's runtime sources:
bash smdd-runtime/build_images.sh oracle
Prime launches task and verifier containers automatically. You start the oracle service yourself.
Ray workers only need the HTTP oracle image.
3. Start the oracle
The oracle exposes the same HTTP API whether it runs on one GPU host or across several GPU VMs with Ray.
Use private or VPN addresses that are mutually reachable. In particular, the oracle must be reachable from both the evaluation host and the verifier containers. 127.0.0.1 inside a verifier points back to that verifier container, not to the evaluation host.
On the GPU host — or on the Ray head — start from the extracted run directory and initialize the persistent state and cache:
export ORACLE_IMAGE="$(python -c 'import json; print(json.load(open("release.json"))["images"]["smdd-oracle-http:latest"])')"
export STATE="$HOME/smdd-oracle/state"
export CACHE="$HOME/smdd-oracle/cache"
install -d -m 700 "$STATE" "$CACHE"
test -s "$STATE/token" || openssl rand -hex 32 > "$STATE/token"
chmod 600 "$STATE/token"
Option A: one GPU host, no Ray
docker run -d --name smdd-oracle-http --restart unless-stopped \
--gpus all --shm-size=16g -p 8090:8090 \
-e ORACLE_ROLE=standalone \
-v "$STATE:/srv/oracle_state" -v "$CACHE:/srv/oracle_cache" \
-v "$STATE/token:/run/secrets/oracle_token:ro" \
"$ORACLE_IMAGE" --max-samples 10 --max-parallel-samples 2 --microbatch-size 1
docker logs --tail 100 -f smdd-oracle-http
This works with one or several GPUs on the same host. docker logs -f only follows the logs: pressing Ctrl+C closes the log viewer and leaves the oracle running.
Option B: two or more GPU VMs with Ray
All nodes need the same oracle image and bidirectional private/VPN connectivity.
Ray uses auxiliary ports in addition to 6379, 6380, 6381, and 20000-20063. Keep these ports on a trusted network. Public NAT addresses are not local bind addresses.
Copy the built oracle image to each worker, replacing the SSH address:
docker save "$ORACLE_IMAGE" | ssh ubuntu@WORKER_SSH_ADDRESS docker load
On the head, use its real private/VPN address and set RAY_EXPECTED_GPUS to the total number of GPUs across the cluster, including the head:
export HEAD_IP="10.77.0.1"
docker run -d --name smdd-head --restart unless-stopped \
--gpus all --network host --shm-size=16g \
-e ORACLE_ROLE=head -e RAY_NODE_IP="$HEAD_IP" \
-e RAY_EXPECTED_GPUS=2 -e ORACLE_HTTP_HOST=0.0.0.0 \
-v "$STATE:/srv/oracle_state" -v "$CACHE:/srv/oracle_cache" \
-v "$STATE/token:/run/secrets/oracle_token:ro" \
"$ORACLE_IMAGE" --max-samples 10 --max-parallel-samples 2 --microbatch-size 1
On each worker, use the same image tag as the head and replace the example address with that worker's own private/VPN address:
export ORACLE_IMAGE="smdd-prime/oracle-http:0.1.4"
export HEAD_IP="10.77.0.1"
export WORKER_IP="10.77.0.2"
docker run -d --name smdd-worker --restart unless-stopped \
--gpus all --network host --shm-size=16g \
-e ORACLE_ROLE=worker -e RAY_NODE_IP="$WORKER_IP" \
-e RAY_ADDRESS="$HEAD_IP:6379" "$ORACLE_IMAGE"
Start the workers before waiting for HTTP readiness. The head waits until the expected GPU actors have initialized.
Workers do not need the oracle bearer token or a shared filesystem. If a node does not appear, inspect Ray and the container logs:
docker exec smdd-head /opt/oracle/bin/ray status
4. Connect Prime to the oracle
There are two pieces of configuration:
- the URL of the running HTTP oracle;
- the local token file containing its authentication token.
These are deployment settings, not fixed defaults in the package.
If the oracle is running on the evaluation VM, inspect its network binding and token mount. For a Ray deployment, use smdd-head in place of smdd-oracle-http.
docker ps --format 'table {{.Names}}\t{{.Status}}\t{{.Ports}}'
export ORACLE_CONTAINER=smdd-oracle-http
docker inspect "$ORACLE_CONTAINER" --format 'Network={{.HostConfig.NetworkMode}} Ports={{json .NetworkSettings.Ports}}'
export SMDD_AGENT_ORACLE_HTTP_TOKEN_FILE="$(docker inspect "$ORACLE_CONTAINER" --format '{{range .Mounts}}{{if eq .Destination "/run/secrets/oracle_token"}}{{.Source}}{{end}}{{end}}')"
printf 'Local token file: %s\n' "$SMDD_AGENT_ORACLE_HTTP_TOKEN_FILE"
The command above prints the path to the token file, not the token itself. If it prints an empty path, check the container name and mounts before continuing.
If the oracle lives on another VM, copy the token securely to the evaluation host and set SMDD_AGENT_ORACLE_HTTP_TOKEN_FILE to that local absolute path.
Choose the oracle URL from the way the service is actually bound:
| Oracle setup | URL to use |
|---|---|
Same-VM mapping 172.17.0.1:8090->8090/tcp | http://172.17.0.1:8090 |
Mapping 0.0.0.0:8090->8090/tcp | The oracle VM's private/VPN IP on port 8090, reachable from the evaluation host and verifier containers |
| Ray head with host networking | The head's reachable private/VPN IP on port 8090; no Docker port mapping is expected |
Use the address and port reported by your deployment. 172.17.0.1 is the Docker bridge address in the tested same-VM setup; it is not a universal oracle address and should not be used from another VM.
For the same-VM bridge setup above, configure the agent and verifier clients like this. Change only the URL if your deployment differs:
export SMDD_AGENT_ORACLE_HTTP_URL="http://172.17.0.1:8090"
export SMDD_AGENT_ORACLE_CACHE_DIR="$HOME/.cache/smdd-prime"
export SMDD_EVALUATOR_ORACLE_HTTP_URL="$SMDD_AGENT_ORACLE_HTTP_URL"
export SMDD_EVALUATOR_ORACLE_HTTP_TOKEN="$(cat "$SMDD_AGENT_ORACLE_HTTP_TOKEN_FILE")"
curl --fail --silent --show-error \
-H "Authorization: Bearer $SMDD_EVALUATOR_ORACLE_HTTP_TOKEN" \
"$SMDD_AGENT_ORACLE_HTTP_URL/v1/health" | python -m json.tool
Initial model loading can take several minutes. Wait until the health endpoint reports healthy model status before starting evaluations.
Keep --max-samples 10 on the oracle. SMDD verifiers request ten diffusion samples regardless of how many GPUs are serving requests.
5. Smoke-test the installation
Before paying for a model rollout, run the supplied reference solutions through the real verifier.
The oracle harness mode does not call a language model. It stages the reference answer and scores it with the same verifier used for a real run. Science tasks still use the oracle.
Start with Lead Optimization, then test one task from every family:
smdd-bench eval @ lead-smoke.toml
smdd-bench eval @ smoke.toml
A healthy installation should finish without runtime exceptions and report:
verifier_error = 0
Do not treat reward = 0 by itself as an infrastructure failure. The scientific reward is task-dependent: a valid molecule can be scored successfully and still fail one of the benchmark gates. If a smoke task returns zero, inspect the individual verifier metrics first.
6. Run a real model trial
The generated lead.toml runs one real Lead Optimization task. For example, to use Claude Sonnet 4.6 through Anthropic:
read -rsp 'Anthropic API key: ' ANTHROPIC_API_KEY
printf '\n'
export ANTHROPIC_API_KEY
smdd-bench eval @ lead.toml \
--model claude-sonnet-4-6 \
--client.base-url https://api.anthropic.com/v1 \
--client.api-key-var ANTHROPIC_API_KEY
Prime routes model calls through its interception endpoint so the full trajectory can be captured.
The original SMDD harness exposes Python, Boltz, ADMET, and submission tools. Its default per-trial budgets are:
- 100 turns
- 8 Boltz calls
- 15 ADMET calls
Change these under [env.agent.harness] if you want a different budget.
Run another task or a larger set
Copy lead.toml and change [env.taskset].tasks to the canonical IDs printed by:
smdd-bench list
Set the top-level num_tasks to the number of task IDs in that list.
Two settings control test-time scaling:
num_rollouts: independent repeats per task;max_concurrent: simultaneous trials.
Start with max_concurrent = 1. Once one end-to-end trial is stable, increase concurrency while watching GPU memory, host memory, and oracle throughput.
7. Follow a run and inspect the results
Prime writes its native traces and results under:
outputs/
The SMDD harness also keeps a per-trial directory under:
smdd-prime-logs/
A completed real trial contains the trajectory, a readable transcript, and — once the agent has terminated — agent_result.json.
The native Prime trace points back to this directory through info.smdd_logs.
To see active trajectories:
find smdd-prime-logs -name trajectory.jsonl
Then follow one of the returned paths:
tail -f smdd-prime-logs/TASK-AND-RUN-ID/trajectory.jsonl
Use a real path returned by find rather than the placeholder above.
A quiet terminal, or a container that is still running, does not necessarily mean model calls are progressing. Check trajectory timestamps and tool responses, then inspect the final verifier metrics.
Completed workers also preserve:
worker.stdout
worker.stderr
The example configs use push = false, so evaluation results are not uploaded automatically.
Deployment boundary
This guide covers Linux + Docker evaluation with an operator-managed GPU oracle.
Hosted sandbox execution is a different deployment mode: task/verifier images must be accessible to the hosted runtime, and the oracle must be reachable from that infrastructure.