TypeSafe Workflow Evals: Security incidents
Frontier
Triage of security incidents as a typed decision, from TypeSafe's four-workflow launch benchmark for Jev.
- Publisher
- TypeSafe AI
- Domain
- knowledge-work
- Published
- Sep 2026
- Notable for
- One of the four workflows Jev was launched on, run against five frontier LLMs.
- Canonical
- evals.typesafe.ai
- Official leaderboard
- evals.typesafe.ai
Cite
Notes
Only stored in your browser.
Top score 66.2% by Claude Opus 5 - 6 models reporting (5 frontier)
Score history
5Top models
6Where it's ranked
1FAQ
- What is TypeSafe Workflow Evals: Security incidents?
- Triage of security incidents as a typed decision, from TypeSafe's four-workflow launch benchmark for Jev.
- What is the current top score on TypeSafe Workflow Evals: Security incidents?
- The top reported score is 66.2% by Claude Opus 5, across 6 models reporting (5 from frontier labs).