Sophon
Catalog of AI evals, the tools that lift them, and the labs behind them. Press ⌘K to search.
Latest
Newest evals, tools, models, and papers
Large language model (LLM) agents can benefit from reusable skills distilled from prior task experience, yet existing skill optimization methods often rely on costly execution-based evaluation and substantial task data.
Reducing audio-encoder depth lowers the inference cost of speech large language models, but removing complete blocks perturbs the embeddings consumed by the decoder and can cause deletion and premature end-of-sequence errors.
Benchmark researchers and developers of large language models (LLMs) and other AI systems need to find relevant evaluations, locate their benchmark datasets and code, and understand the settings behind reported scores.
On-Policy Self-Distillation (OPSD) has emerged as a popular paradigm for large language model (LLM) self-improvement, allowing models to act as their own teachers by leveraging privileged information such as ground-truth solutions.
We study how model post-training and test-time inference design affect natural-language proof generation for hard olympiad mathematics. Starting from Nemotron 3 Ultra, we train two specialist checkpoints using supervised fine-tuning and reinforcement learning, and evaluate…
Multimodal entity linking grounds entity mentions in text and images to knowledge-base entries. These systems degrade on rare entities, but prior work measures rarity primarily through popularity-based metrics such as pageviews.
We introduce NCP-ArchPreview, a latent-space language model that pushes autoregressive pretraining beyond standard next-token prediction (NTP). Alongside NTP, the model learns through Next Concept Prediction (NCP) to predict discrete concepts that span multiple tokens,…
Leaderboards
Current standings across rating systems
#1 Claude Opus 5 · 1504
#1 Claude Fable 5 · 1507
#1 Claude Opus 4.6 · 1523
#1 Claude Opus 4.6 · 1527
#1 Claude Fable 5 · 1506
#1 Claude Opus 4.6 · 1504
What lifts scores most
Tools with the most known eval-lift evidence
Browse by capability
Pick a capability to rank the tools that train toward it
Closest to saturation
Benchmarks where the top score approaches the ceiling

