0

ARC-AGI-3

Saturated

Interactive generalisation benchmark from the ARC Prize Foundation, run as an agent over unfamiliar environments. Scores depend heavily on the surrounding harness, not the model alone.

Domain
agent-eval
Published
Sep 2026
Notable for
The generalisation benchmark behind the AGI framing of the GPT-6 Astra launch.

Cite

Notes

Only stored in your browser.

Attribution

Leaderboard scores
OpenAI
Attribution policy →

Top score 99.9% by GPT-6 Astra - 1 model reporting (1 frontier)

Top models

1
ARC-AGI-3Bar chart with 1 bar. Highest value: GPT-6 Astra at 99.9.
1 model

FAQ

What is ARC-AGI-3?
Interactive generalisation benchmark from the ARC Prize Foundation, run as an agent over unfamiliar environments. Scores depend heavily on the surrounding harness, not the model alone.
What is the current top score on ARC-AGI-3?
The top reported score is 99.9% by GPT-6 Astra, across 1 model reporting (1 from frontier labs).