ARC-AGI-3
Saturated
Interactive generalisation benchmark from the ARC Prize Foundation, run as an agent over unfamiliar environments. Scores depend heavily on the surrounding harness, not the model alone.
- Domain
- agent-eval
- Published
- Sep 2026
- Notable for
- The generalisation benchmark behind the AGI framing of the GPT-6 Astra launch.
Cite
Notes
Only stored in your browser.
Top score 99.9% by GPT-6 Astra - 1 model reporting (1 frontier)
Top models
1FAQ
- What is ARC-AGI-3?
- Interactive generalisation benchmark from the ARC Prize Foundation, run as an agent over unfamiliar environments. Scores depend heavily on the surrounding harness, not the model alone.
- What is the current top score on ARC-AGI-3?
- The top reported score is 99.9% by GPT-6 Astra, across 1 model reporting (1 from frontier labs).