Artificial Analysis - Intelligence Index
Aggregate Intelligence Index (0-100) over MMLU-Pro, GPQA-Diamond, HumanEval, MATH-500, and other reasoning benchmarks. Published by Artificial Analysis with per-model pricing, throughput, and latency.
- Operator
- Artificial Analysis
- Kind
- Aggregated
- Updates
- weekly·updated 17m ago
- Notable for
- intelligence-index
- Tracks
- 12 evals · aggregated
Cite
Notes
Only stored in your browser.
Intelligence ranking
Per-eval breakdown
485models
| Model | |||||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| R1 1776 Perplexity AI | - | - | - | - | - | - | 95.4% | - | - | - | - | - | 95.4% |
| o1 Preview OpenAI | - | - | - | - | - | - | 92.4% | - | - | - | - | - | 92.4% |
| o3 Pro OpenAI | - | - | 84.5% | - | - | - | - | - | - | - | - | - | 84.5% |
| Hermes 4 (405B) Nous Research | 81.9% | - | - | - | - | - | - | - | - | - | - | - | 81.9% |
| DeepSeek-V2.5 (Dec '24) DeepSeek | - | - | - | - | - | - | 76.3% | - | - | - | - | - | 76.3% |
| DeepSeek-Coder-V2 DeepSeek | - | - | - | - | - | - | 74.3% | - | - | - | - | - | 74.3% |
| Gemini 3 Pro Google (Alphabet Inc.) | - | 95.7% | 91.9% | 37.5% | 70.4% | 91.7% | - | 89.8% | 56.1% | 87.1% | 41.7% | - | 73.5% |
| Gemini 3 Flash Preview Google (Alphabet Inc.) | - | 97.0% | 89.8% | 36.6% | 78.0% | 90.8% | - | 89.0% | 50.6% | 80.4% | 38.6% | - | 72.3% |
| Claude Fable 5 Anthropic | - | - | 92.6% | 55.5% | 63.5% | - | - | - | 60.2% | 98.5% | 62.9% | - | 72.2% |
| o3 OpenAI | 96.7% | 88.3% | 87.7% | 20.1% | 71.4% | 80.8% | 99.2% | 85.3% | 41.0% | 80.7% | 37.1% | - | 71.7% |
| Gemini 3.1 Pro Preview Google (Alphabet Inc.) | - | - | 94.1% | 47.0% | 77.1% | - | - | - | 58.9% | 95.6% | 53.8% | - | 71.1% |
| Grok 4 xAI | 94.3% | 92.7% | 87.7% | 26.7% | 53.7% | 81.9% | 99.0% | 86.6% | 45.7% | 74.9% | 37.9% | - | 71.0% |
| Gemini 2.5 Pro Preview (Mar' 25) Google (Alphabet Inc.) | 87.0% | - | 83.6% | 18.0% | - | 77.8% | 98.0% | 85.8% | 39.5% | - | - | - | 70.0% |
| Gemini 2.5 Pro Preview (May' 25) Google (Alphabet Inc.) | 84.3% | - | 82.2% | 19.9% | - | 77.0% | 98.6% | 83.7% | 41.6% | - | - | - | 69.6% |
| GPT-5 Codex (batch) OpenAI | - | 98.7% | 83.7% | 27.8% | 74.1% | 84.0% | - | 86.5% | 40.9% | 86.8% | 37.9% | - | 68.9% |
| Claude Opus 4.8 Anthropic | - | - | 92.0% | 48.7% | 62.2% | - | - | - | 53.5% | 94.4% | 58.3% | - | 68.2% |
| Qwen3.7 Max Alibaba | - | - | 92.3% | 40.5% | 80.5% | - | - | - | 48.8% | 94.7% | 50.8% | - | 67.9% |
| MiMo-V2.5-Pro Xiaomi | - | - | 86.6% | 35.7% | 79.9% | - | - | 85.1% | 50.2% | 94.2% | 43.2% | - | 67.8% |
| Kimi K2 Thinking Kimi | - | 94.7% | 83.8% | 23.8% | 68.1% | 85.3% | - | 84.8% | 42.4% | 93.0% | 31.1% | - | 67.4% |
| Gemini 3 Deep Think Google DeepMind | - | - | 93.8% | 41.0% | - | - | - | - | - | - | - | - | 67.4% |
| GPT-5.1-Codex OpenAI | - | 95.7% | 86.0% | 25.7% | 70.0% | 84.9% | - | 86.0% | 40.2% | 83.0% | 34.8% | - | 67.4% |
| GPT-5.3-Codex OpenAI | - | - | 91.5% | 42.5% | 75.4% | - | - | - | 53.2% | 86.0% | 53.0% | - | 66.9% |
| o4 Mini OpenAI | 94.0% | 90.7% | 78.4% | 16.5% | 68.7% | 85.9% | 98.9% | 83.2% | 46.5% | 55.6% | 15.2% | - | 66.7% |
| Kimi K3 Kimi | - | - | 93.5% | 46.9% | - | - | - | - | 58.7% | - | - | - | 66.4% |
| Kimi K2.6 Moonshot AI | - | - | 91.1% | 37.5% | 76.0% | - | - | - | 53.5% | 95.9% | 43.9% | - | 66.3% |
| Muse Spark Meta Platforms | - | - | 88.4% | 40.7% | 75.9% | - | - | - | 51.5% | 91.5% | 45.5% | - | 65.6% |
| Gemini 2.5 Pro Google (Alphabet Inc.) | 88.7% | 87.7% | 84.4% | 22.5% | 48.7% | 80.1% | 96.7% | 86.2% | 42.8% | 54.1% | 26.5% | - | 65.3% |
| MiniMax M3 Minimax | - | - | 92.9% | 39.0% | 82.9% | - | - | - | 45.4% | 88.9% | 42.4% | - | 65.2% |
| Grok 3 mini xAI | 93.3% | 84.7% | 79.1% | 11.0% | 45.9% | 69.6% | 99.2% | 82.8% | 40.6% | 90.4% | 17.4% | - | 64.9% |
| Qwen3.7 Plus Alibaba | - | - | 90.0% | 35.6% | 78.0% | - | - | - | 45.5% | 93.0% | 47.0% | - | 64.8% |
| Muse Spark 1.1 Meta Platforms | - | - | 89.8% | 46.2% | - | - | - | - | 58.2% | - | - | - | 64.7% |
| Claude Mythos Preview Anthropic | - | - | - | 64.7% | - | - | - | - | - | - | - | - | 64.7% |
| MiniMax M2.1 Minimax | - | 82.7% | 83.0% | 23.2% | 69.9% | 81.0% | - | 87.5% | 40.7% | 85.4% | 28.8% | - | 64.7% |
| Gemini 3 Pro Preview Google (Alphabet Inc.) | - | 86.7% | 88.7% | 29.5% | 49.7% | 85.7% | - | 89.5% | 49.9% | 68.1% | 34.1% | - | 64.6% |
| GPT-5.2-Codex OpenAI | - | - | 89.9% | 35.7% | 77.6% | - | - | - | 54.6% | 92.1% | 37.1% | - | 64.5% |
| Muse Spark 1.2 Meta Platforms | - | - | 90.4% | 45.5% | - | - | - | - | 56.4% | - | - | - | 64.1% |
| Qwen3.6 Max Preview Alibaba | - | - | 88.8% | 30.8% | 76.6% | - | - | - | 46.9% | 95.9% | 43.9% | - | 63.8% |
| Qwen3 235B A22B Thinking 2507 Alibaba | 94.0% | 91.0% | 79.0% | 15.9% | 51.2% | 78.8% | 98.4% | 84.3% | 42.4% | 53.2% | 13.6% | - | 63.8% |
| Grok 4.5 xAI | - | - | 93.1% | 42.7% | - | - | - | - | 54.1% | - | - | - | 63.3% |
| Qwen3.8 Max Alibaba | - | - | 92.7% | 43.0% | - | - | - | - | 52.9% | - | - | - | 62.9% |
| GPT-5.1-Codex-Mini OpenAI | - | 91.7% | 81.3% | 18.5% | 67.9% | 83.6% | - | 82.0% | 42.6% | 62.9% | 33.3% | - | 62.6% |
| KAT-Coder-Pro V1 KwaiKAT | - | 94.7% | 76.4% | 33.6% | 68.4% | 74.7% | - | 81.3% | 36.6% | 88.6% | 9.1% | - | 62.6% |
| Nex-N2-Pro Nex AGI | - | - | 89.2% | 33.7% | 66.2% | - | - | - | 41.8% | 81.6% | - | - | 62.5% |
| Qwen3.6 Plus Alibaba | - | - | 88.2% | 27.8% | 75.2% | - | - | - | 40.7% | 97.7% | 43.9% | - | 62.2% |
| Gemini 3.6 Flash Google (Alphabet Inc.) | - | - | 92.8% | 40.8% | - | - | - | - | 52.7% | - | - | - | 62.1% |
| MiniMax M2 Minimax | - | 78.3% | 77.7% | 13.7% | 72.3% | 82.6% | - | 82.0% | 36.1% | 86.8% | 25.8% | - | 61.7% |
| Kimi K2.7 Code Kimi | - | - | 89.6% | 35.0% | 63.1% | - | - | - | 47.5% | 90.1% | 44.7% | - | 61.7% |
| Gemini 2.5 Flash Preview (Reasoning) Google (Alphabet Inc.) | 84.3% | - | 69.8% | 12.1% | - | 50.5% | 98.1% | 80.0% | 35.9% | - | - | - | 61.5% |
| MiMo-V2-Pro Xiaomi | - | - | 87.0% | 30.4% | 68.8% | - | - | - | 42.5% | 95.0% | 40.9% | - | 60.8% |
| MiniMax M2.7 Minimax | - | - | 87.4% | 29.6% | 75.7% | - | - | - | 47.0% | 84.8% | 39.4% | - | 60.7% |
