0

Reliability Scales Inversely: Bigger Models Compound Mistakes Faster via a Hidden Auto-Regressive Risk Regime

As language models scale, answers start truer but degrade faster: scaling buys capability but erodes reliability. The knowledge-gap account - more data, retrieval, or scale - misses an auto-regressive risk residual that scale sharpens: the model commits to a low-probability…

Preview
Year
2026
Hosting
Full text hostedCC-BY-4.0

Cite

Notes

Only stored in your browser.

Attribution

Abstract & full text
arxiv.org/abs/2607.18292CC-BY-4.0
TL;DR
Semantic Scholar
Attribution policy →

Abstract

As language models scale, answers start truer but degrade faster: scaling buys capability but erodes reliability. The knowledge-gap account - more data, retrieval, or scale - misses an auto-regressive risk residual that scale sharpens: the model commits to a low-probability token, conditions on it as established, and snowballs. We track this through per-position disagreement δ= \log p_M - \log p_O against a stronger same-family oracle, whose second moment splits exactly into bias^2 KL(p_M ,|, p_O)^2 and risk Var[δ]. We present four findings: (i) under scaling, the knowledge gap falls \approx6\times while knowledge degradation grows 11-39\times; (ii) at a fabrication, felt uncertainty H(p_M) relaxes quickly while oracle-referenced risk persists up to 17\times longer, leaving a confident-but-precarious risk regime that bridges consecutive fabrications (+69% at 14B); (iii) this regime is causal - an on-policy, fixed-KL variance contraction cuts web-verified hallucination by 35-74% across three model families; and, (iv) it structurally evades self-monitoring, with p_M-only detectors (e.g. semantic entropy) firing \approx30% less (p<10^{-16}) on the risky branch holding nearly 4\times more fabrications. Bigger models snowball mistakes faster, through a failure mode that is dominant, self-perpetuating, causal and invisible to the model itself.