0

Sophon

Catalog of AI evals, the tools that lift them, and the labs behind them.

Frontier right now

Top models on Artificial Analysis - Intelligence Index

Artificial Analysis - Intelligence IndexBar chart with 20 bars. Highest value: Claude Fable 5.1 at 53.4.
20 models

Latest

View all

Newest evals, tools, models, and papers

Model4d ago
Agnes 3.0 Flash
Sapiens AI
$0.07/M235 tok/s
GPQA Diamond92
SciCode52
Humanity's Last Exam (HLE)39
Model5d ago
DeepSeek V4.1 Flash
DeepSeek
1.0M$0.53/M227 tok/smitOpen
MedScribe86
Vibe Code Bench v1.185
Vals Index58
Model5d ago
Ling 3.0 Flash VL
InclusionAI
262K147 tok/smitOpen
GPQA Diamond86
SciCode44
Humanity's Last Exam (HLE)22
Paper5d ago
COBRA-Skills: Contextual Bandit-Guided Evolution for Agent Skill Optimization

Large language model (LLM) agents can benefit from reusable skills distilled from prior task experience, yet existing skill optimization methods often rely on costly execution-based evaluation and substantial task data.

Agents
8stars
0.1/h
Paper5d ago
X-AuT: Progressive Audio-Encoder Compression for Speech LLMs with Cross-Scale Distillation

Reducing audio-encoder depth lowers the inference cost of speech large language models, but removing complete blocks perturbs the embeddings consumed by the decoder and can cause deletion and premature end-of-sequence errors.

Automatic Speech Recognition
10stars
0.0/h
Paper5d ago
Benchmark Radar: A Living Database and Search Engine for AI Benchmarks and Evaluation

Benchmark researchers and developers of large language models (LLMs) and other AI systems need to find relevant evaluations, locate their benchmark datasets and code, and understand the settings behind reported scores.

Retrieval
204stars
0.7/h
Paper5d ago
Negative Self-Distillation: Learning to Reason by Avoiding Flaws

On-Policy Self-Distillation (OPSD) has emerged as a popular paradigm for large language model (LLM) self-improvement, allowing models to act as their own teachers by leveraging privileged information such as ground-truth solutions.

ReasoningReinforcement Learning
8stars
0.1/h
Paper6d ago
An Open Recipe for IMO Gold: Training Nemotron for Olympiad Mathematics

We study how model post-training and test-time inference design affect natural-language proof generation for hard olympiad mathematics. Starting from Nemotron 3 Ultra, we train two specialist checkpoints using supervised fine-tuning and reinforcement learning, and evaluate…

MathematicsReasoningReinforcement Learning
1.0kstars
0.1/h
Paper6d ago
Think Before You Link: Rarity, Reasoning, and Retrieval in Multilingual Entity Linking

Multimodal entity linking grounds entity mentions in text and images to knowledge-base entries. These systems degrade on rare entities, but prior work measures rarity primarily through popularity-based metrics such as pageviews.

Document UnderstandingReasoningRetrieval
3stars
0.0/h
Paper6d ago
NCP-ArchPreview Technical Report: Moving towards Latent Space Language Models through Next Concept Prediction

We introduce NCP-ArchPreview, a latent-space language model that pushes autoregressive pretraining beyond standard next-token prediction (NTP). Alongside NTP, the model learns through Next Concept Prediction (NCP) to predict discrete concepts that span multiple tokens,…

Language ModelingMathematics
17stars
0.2/h

Leaderboards

View all

Current standings across rating systems

What lifts scores most

View all

Tools with the most known eval-lift evidence

Browse by capability

View all

Pick a capability to rank the tools that train toward it

Closest to saturation

View all

Benchmarks where the top score approaches the ceiling