Agent-safety benchmarks measure different behaviors, and their scores get quoted interchangeably as an agent's safety. We treat four of them (R-Judge, InjecAgent, AgentHarm, AgentDojo) as measurements to be validated, running each under its official implementation and author-provided scorer on up to 22 models, with MMLU and GPQA measured by us under one protocol as a capability composite. The metric is the first problem. On any binary trace-judgment benchmark scored by F_1, an ``always positive'' policy attains F_1 = 2π/(1+π); on R-Judge that is 0.690, above five of the 21 models that actually discriminate. The three broad-coverage benchmarks then rank the same 18 models differently, and the trade-off behind that disagreement is a small-panel artifact: R-Judge specificity against AgentHarm safety correlates -0.64 at n{=}7 and +0.02 at n{=}18, and a quarter of random size-7 subsets reach |ρ| \geq 0.5 around that near-zero value. Held-out validity turns on which outcome you pick. Capability predicts task success (ρ{=}{+}0.60) but correlates negatively with misalignment safety (ρ{=}{-}0.44, n{=}21). On their paired n{=}20 panel, the corresponding contrast is Δ{=}{-}1.00 (95% CI [-1.48, -0.49], p<0.001), and it survives leave-one-organization-out and organization-clustered bootstrap analyses. On an expanded 41-model panel, the misalignment correlation weakens to -0.16 (95% CI [-0.54, +0.22]) and jailbreak strengthens to +0.34, though neither change is significant. \mbox{AgentHarm} shows the strongest held-out association, ρ{=}{+}0.72 with three-template jailbreak safety after controlling capability. But both instruments score harmful compliance, so this is evidence of convergent validity rather than general safety. Naming the benchmark, metric, target behavior, and model panel is the minimum a safety claim needs.
Safety, or Just Capability? A Validity Audit of Agent-Safety Benchmarks
Agent-safety benchmarks measure different behaviors, and their scores get quoted interchangeably as an agent's safety. We treat four of them (R-Judge, InjecAgent, AgentHarm, AgentDojo) as measurements to be validated, running each under its official implementation and…
- Preview

- Year
- 2026
- Hosting
- Full text hostedCC-BY-4.0
Cite
Notes
Only stored in your browser.
Attribution
- Abstract & full text
- arxiv.org/abs/2607.28685CC-BY-4.0
- TL;DR
- Semantic Scholar