0

A Reproducibility Study of Partial Residual Ablations in Pre-LN Transformers

Residual connections are a fundamental component of transformer architectures, yet the roles of the attention and feed-forward residual pathways remain poorly understood when considered independently.

Preview
Year
2026
Hosting
Full text hostedCC-BY-4.0

Cite

Notes

Only stored in your browser.

Attribution

Abstract & full text
arxiv.org/abs/2608.14689CC-BY-4.0
TL;DR
Semantic Scholar
Attribution policy →

Abstract

Residual connections are a fundamental component of transformer architectures, yet the roles of the attention and feed-forward residual pathways remain poorly understood when considered independently. This paper presents a reproducibility study of partial residual ablations in Pre-LN GPT-style transformers trained at two scales (10M and 124M parameters). I compare four architectural configurations by selectively removing the attention residual connection, the feed-forward residual connection, or both. At 10M scale, a controlled 8-seed deterministic sweep shows a clear asymmetry: removing the attention residual (FFNOnly) reaches the No-Residual collapse regime (3.350 +/- 0.002), whereas removing the feed-forward residual (AttnOnly) remains well separated from it (1.580 +/- 0.003). A subsequent controlled 124M sweep establishes that the direction of this asymmetry persists under matched seeds: across five seeds each, AttnOnly reaches 6.037 +/- 0.355 versus 7.569 +/- 0.001 for FFNOnly, with no overlap between the observed ranges. AttnOnly nevertheless remains substantially degraded relative to FullResidual (4.577 +/- 0.012) and exhibits markedly greater seed sensitivity at 124M. During the investigation, I identified and corrected an experimental measurement confound in runtime gain scaling and retained an intermediate reproduction failure rather than excluding it. I propose cross-position routing through self-attention as a falsifiable hypothesis for the observed asymmetry; no clean causal test of this mechanism has yet been completed. To support reproducibility, I release the source code, experiment configurations, training logs, and experimental results, including intermediate non-reproducing runs.