0

Dimensionality and Measurement Precision in HLE's Multiple-Choice Subset

Humanity's Last Exam (HLE) is widely used to evaluate frontier language models. HLE organizes its questions into eight subject-domain categories, whose subscores are often interpreted as evidence of distinct capabilities.

Preview
Year
2026
Hosting
Full text hostedCC-BY-4.0

Cite

Notes

Only stored in your browser.

Attribution

Abstract & full text
arxiv.org/abs/2607.27420CC-BY-4.0
TL;DR
Semantic Scholar
Attribution policy →

Abstract

Humanity's Last Exam (HLE) is widely used to evaluate frontier language models. HLE organizes its questions into eight subject-domain categories, whose subscores are often interpreted as evidence of distinct capabilities. However, no study has assessed whether these labels correspond to empirically separable latent constructs, nor whether the benchmark effectively differentiates between models of similar ability. We evaluate 29 LLMs on the text-only multiple-choice subset of HLE (J = 428 items) and apply psychometric methods to assess both the dimensionality of the benchmark and the distribution of its measurement precision. Fitting a two-parameter logistic IRT model, we find convergent evidence that HLE measures a single general reasoning factor: McDonald's ω_h = 0.998, domain labels explain only 3.5% of item response variance, within- and between-domain residual correlations are nearly identical (Cohen's d = 0.016), and domain-specific ability estimates are near-redundant with the total score (r \geq 0.81). A separate analysis of the test information function reveals that measurement precision concentrates at moderate ability levels and drops sharply above θ= 0, where frontier models sit. These findings suggest that HLE's domain subscores do not warrant distinct capability interpretations and that the benchmark's ability to discriminate among the strongest models is limited.