0

TypeSafe Workflow Evals: Agent trace observability

Frontier

Judging agent execution traces as a typed decision, from TypeSafe's four-workflow launch benchmark for Jev.

Publisher
TypeSafe AI
Domain
knowledge-work
Published
Sep 2026
Notable for
One of the four workflows Jev was launched on, run against five frontier LLMs.
Official leaderboard
evals.typesafe.ai

Cite

Notes

Only stored in your browser.

Attribution

Leaderboard scores
typesafe-release
Attribution policy →

Top score 76.1% by GPT-5.6 Luna - 6 models reporting (5 frontier)

Score history

5
60%70%80%90%100%Jun 26Jul 26Aug 26Sep 26GPT-5.6 Luna

Top models

6
TypeSafe Workflow Evals: Agent trace observabilityBar chart with 6 bars. Highest value: GPT-5.6 Luna at 76.1.
6 models

Where it's ranked

1

FAQ

What is TypeSafe Workflow Evals: Agent trace observability?
Judging agent execution traces as a typed decision, from TypeSafe's four-workflow launch benchmark for Jev.
What is the current top score on TypeSafe Workflow Evals: Agent trace observability?
The top reported score is 76.1% by GPT-5.6 Luna, across 6 models reporting (5 from frontier labs).