TypeSafe Workflow Evals: Agent trace observability
Frontier
Judging agent execution traces as a typed decision, from TypeSafe's four-workflow launch benchmark for Jev.
- Publisher
- TypeSafe AI
- Domain
- knowledge-work
- Published
- Sep 2026
- Notable for
- One of the four workflows Jev was launched on, run against five frontier LLMs.
- Canonical
- evals.typesafe.ai
- Official leaderboard
- evals.typesafe.ai
Cite
Notes
Only stored in your browser.
Top score 76.1% by GPT-5.6 Luna - 6 models reporting (5 frontier)
Score history
5Top models
6Where it's ranked
1FAQ
- What is TypeSafe Workflow Evals: Agent trace observability?
- Judging agent execution traces as a typed decision, from TypeSafe's four-workflow launch benchmark for Jev.
- What is the current top score on TypeSafe Workflow Evals: Agent trace observability?
- The top reported score is 76.1% by GPT-5.6 Luna, across 6 models reporting (5 from frontier labs).