Sophon
Catalog of AI evals, the tools that lift them, and the labs behind them. Press ⌘K to search.
Leaderboards
Current standings across rating systems
Arena - TextLMArena
#1 Claude Opus 5 · 1504
Arena - Text Style ControlLMArena
#1 Claude Fable 5 · 1507
Arena Hard PromptsLMArena
#1 Claude Opus 4.6 · 1527
LMArena - Instruction FollowingLMArena
#1 Claude Opus 4.6 · 1523
LMArena - Writing and LiteratureLMArena
#1 Claude Fable 5 · 1506
LMArena - Creative WritingLMArena
#1 Claude Opus 4.6 · 1504
What lifts scores most
Tools with the most known eval-lift evidence
Browse by capability
Pick a capability to rank the tools that train toward it
Closest to saturation
Benchmarks where the top score approaches the ceiling
Mostly Basic Python Problems (MBPP)15 models100.0%MATH-500190 models99.4%τ²-bench (Tau²-bench)321 models99.1%AIME 2025: Problems from the American Invitational Mathematics Examination205 models99.0%AIME 2024: Problems from the American Invitational Mathematics Examination169 models96.7%LiveBench - Math59 models96.3%GPQA Diamond501 models95.3%GPQA (Full Set)10 models93.2%
