Thematic Generalization Benchmark
Frontier
Given a group of items sharing a hidden theme plus deliberate 'anti-examples', the model must infer the theme and rank candidates by fit. Scored by how high the true answer lands. Published by Lech Mazur.
- Domain
- Reasoning
- Format
- Custom
- Published
- Jan 2025
- Canonical
- github.com/lechmazur/generalization
- Official leaderboard
- github.com/lechmazur/generalization
Cite
Notes
Only stored in your browser.
Top score 80.6 by Claude Opus 4.6 - 21 models reporting (7 frontier)
Score history
20Top models
21Where it's ranked
1FAQ
- What is Thematic Generalization Benchmark?
- Given a group of items sharing a hidden theme plus deliberate 'anti-examples', the model must infer the theme and rank candidates by fit. Scored by how high the true answer lands. Published by Lech Mazur.
- What is the current top score on Thematic Generalization Benchmark?
- The top reported score is 80.6 by Claude Opus 4.6, across 21 models reporting (7 from frontier labs).

