Claude Opus 4.8 vs GPT-4.1 Mini
Head to head on 11 shared benchmarks, with price, context window, and release dates.
Benchmark scores
170evals · sort by any column, ⤢ to expand fullscreen
| Model | |||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Claude Opus 4.8 Anthropic | 66.7% | 1835elo | 47.2% | 92.0% | 48.7% | 62.2% | 69.9% | 53.5% | 75.6% | 58.3% | 94.4% | - | - | - | - | - | - | - | - | - | - | - | - | - | - | - | 15.5% | - | - | - | 80.4% | - | - | 14.5% | - | - | - | - | - | - | - | - | - | - | - | - | - | 67.0% | - | - | - | - | - | - | - | 1281elo | 80.8 | 40.0% | - | - | 53.9% | - | - | - | 13.4% | - | - | 22.5% | 1890elo | - | - | - | - | - | - | - | - | 56.9% | - | - | - | - | - | - | - | - | 10.4% | - | - | - | 80.1% | 79.3% | 78.3% | 67.4% | 81.4% | 84.3% | 89.7% | - | - | - | - | - | - | - | - | 91.8% | - | - | - | 53.2% | - | 85.8% | - | - | - | - | - | - | - | - | - | - | - | - | - | - | 83.4% | - | - | 69.0% | - | - | - | - | - | 54.8% | - | - | - | - | - | - | - | - | - | - | - | - | - | 69.2% | - | - | 82.7% | - | - | - | - | - | - | - | 60.9% | 70.9% | - | - | 82.7% | - | - | - | - | - | 63.4% |
| GPT-4.1 Mini OpenAI | 57.9% | 1144elo | 4.5% | 66.4% | 5.0% | 38.3% | 65.5% | 40.4% | 71.9% | 7.6% | 52.9% | -0.17 | 60.0% | 100.0% | 0.0% | 3 | 32.4% | 43.0% | 46.3% | 98.2% | 0.0% | 28.2% | -1 | 1.07 | 62.3% | 89.9% | - | 100 | 83.62 | 0.79 | - | 0.0% | -0.04 | - | 0.0% | 0.0% | 1.11 | 1.11 | 1.01 | 204.5 | 40.0% | 41.7% | 1.45 | 34.5% | 13.0% | 53.6% | 3.6% | - | 17.4% | 3.3% | 9.0% | 80.0% | 80.0% | 83.4% | 48.6% | - | - | - | 43.3% | 100.0% | - | 58.2% | 52.9% | 0.0% | - | 1.5% | 18.7% | - | - | 50.0% | 65.0% | 56.8% | 0.0% | 100.0% | 3.47 | 3.47 | 1.55 | - | 100.0% | 63.3% | 0.0% | 0.0% | 67.5% | 40.2% | 83.5% | 6.7% | - | 62.4% | 43.8% | 26.7% | - | - | - | - | - | - | - | 48.3% | 45.4% | 80.0% | 0.35 | 100.0% | 93.3% | 100.0% | 92.5% | - | 2.61 | 0.0% | 100.0% | - | 77.9% | - | 5 | 0.0% | 64.0% | 64.0% | 20.0% | 26.7% | 78.1% | 54.5% | 20.0% | 20.0% | 30.0% | 19.5% | 96.5% | 52.0% | - | 30.0% | 30.0% | - | 60.0% | 0.0% | 46.8% | 93.0% | 40.6% | - | 100.0% | 88.2% | 100.0% | 20.0% | 33.3% | 100.0% | 73.3% | 100.0% | 1.14 | 56.72 | 0 | 10.0% | 23.9% | - | 100.0% | 100.0% | - | 0.0% | 100.0% | 37.6% | 1.04 | 100.0% | 100.0% | 45.6% | - | - | 93.3% | 100.0% | - | 29.0% | 90.0% | 29.0% | 82.0% | 82.0% | 50.8% |
2 / 2 models
| Attribute | Claude Opus 4.8 Anthropic | GPT-4.1 Mini OpenAI |
|---|---|---|
| Description | Claude Opus 4.8 (Adaptive Reasoning, Max Effort) is an AI model from Anthropic. | GPT-4.1 Mini is an AI model from OpenAI. |
| Family | Claude | GPT |
| Params | - | - |
| Open weights | - | closed |
| License | - | proprietary |
| Released | 28 May 2026 | 14 Apr 2025 |
| Context | 1,000,000 | 1,047,576 |
| Input $/1M | $5.00 | $0.40 |
| Output $/1M | $25.00 | $1.60 |
| Cache read $/1M | $0.500 | $0.100 |
| Cache write $/1M | $6.250 | - |
| Speed | 0 tok/s | 0 tok/s |
| Latency | 0.00 s | 0.00 s |
| Modalities | text, image, file | image, text, file |
| Providers | - | - |
| Publisher | Anthropic | OpenAI |
| Reported scores |
FAQ
- Is Claude Opus 4.8 or GPT-4.1 Mini better?
- Across the 11 benchmarks both models report on Sophon, Claude Opus 4.8 leads on 11 of them. Which one is "better" depends on the benchmark - the table above shows every shared score.
- How do Claude Opus 4.8 and GPT-4.1 Mini compare on GPQA Diamond?
- Claude Opus 4.8 scores 92.0% and GPT-4.1 Mini scores 66.4% on GPQA Diamond.
- Which is cheaper, Claude Opus 4.8 or GPT-4.1 Mini?
- GPT-4.1 Mini is cheaper at $1.60 per million output tokens against $25.00 for Claude Opus 4.8 - about 15.6x.
- When were Claude Opus 4.8 and GPT-4.1 Mini released?
- Claude Opus 4.8 was released 28 May 2026 by Anthropic. GPT-4.1 Mini was released 14 Apr 2025 by OpenAI.
Comparing something else? Build your own side-by-side.