aime25
AIME 2025 competition math tasks (parts I and II). The model reasons through each task and boxes an integer answer, scored by math equivalence (math-verify) of the boxed answer against the gold.
Taskset
- Source:
opencompass/AIME2025 - Size: 30 tasks (AIME2025-I and AIME2025-II combined)
Changelog
- 2026-08-31: Yield task records on demand so bounded evaluations construct only the requested prefix.
- 2026-06-23: Initial v1 taskset.