Autoregressive transformer language models frequently exhibit training instability when trained on long sequences, particularly under low-precision arithmetic. Although this instability is often accompanied by attention-logit explosion, its underlying cause remains poorly understood. In this work, we present analytical insights and empirical evidence that dense local dependencies are a major contributor to attention-logit explosion. We demonstrate that dense local dependency patterns yield an effectively high-rank attention structure, which the low-rank parameterization of self-attention can only approximate with increasingly large logits as the sequence length grows. This logit inflation ultimately leads to training instability under low-precision arithmetic. We support this explanation through extensive experiments on synthetic and language modeling tasks. Our results consistently show that attention-logit growth increases with sequence length, is mitigated by increasing the attention dimension, and is substantially reduced by explicitly modeling dense local dependencies. Furthermore, we show that this growth is driven by the density of local dependencies rather than by locality alone. More broadly, our findings suggest that explicitly modeling dense local dependencies constitutes an important design principle for developing stable, efficient, and scalable long-context transformer architectures for autoregressive language modeling.
Dense Local Dependencies Induce Attention-Logit Explosion and Training Instability During Long-Sequence Transformer Training
Autoregressive transformer language models frequently exhibit training instability when trained on long sequences, particularly under low-precision arithmetic. Although this instability is often accompanied by attention-logit explosion, its underlying cause remains poorly…
- Year
- 2025
- Hosting
- Full text hostedCC-BY-4.0
Cite
Notes
Only stored in your browser.
Attribution
- Abstract & full text
- arxiv.org/abs/2505.15548CC-BY-4.0
- TL;DR
- Semantic Scholar