0

TypeSafe Workflow Evals: Security incidents

Frontier

Triage of security incidents as a typed decision, from TypeSafe's four-workflow launch benchmark for Jev.

Publisher
TypeSafe AI
Domain
knowledge-work
Published
Sep 2026
Notable for
One of the four workflows Jev was launched on, run against five frontier LLMs.
Official leaderboard
evals.typesafe.ai

Cite

Notes

Only stored in your browser.

Attribution

Leaderboard scores
typesafe-release
Attribution policy →

Top score 66.2% by Claude Opus 5 - 6 models reporting (5 frontier)

Score history

5
45%59%73%86%100%Jun 26Jul 26Aug 26Sep 26GPT-5.6 LunaClaude Opus 5

Top models

6
TypeSafe Workflow Evals: Security incidentsBar chart with 6 bars. Highest value: Claude Opus 5 at 66.2.
6 models

Where it's ranked

1

FAQ

What is TypeSafe Workflow Evals: Security incidents?
Triage of security incidents as a typed decision, from TypeSafe's four-workflow launch benchmark for Jev.
What is the current top score on TypeSafe Workflow Evals: Security incidents?
The top reported score is 66.2% by Claude Opus 5, across 6 models reporting (5 from frontier labs).