0

$σ$Transfer: Uncertainty Transfer from Small to Large Networks under $μ\mathrm{P}$

Reliable predictive uncertainty in Laplace approximations depends critically on the prior precision, yet selecting it requires a posterior sweep that is prohibitively expensive for neural networks with billions of parameters.

Preview
Year
2026
Hosting
Abstract onlyARXIV-DEFAULT

Cite

Notes

Only stored in your browser.

Attribution

Abstract & full text
arxiv.org/abs/2610.11668ARXIV-DEFAULT
TL;DR
Semantic Scholar
Attribution policy →

Abstract

Reliable predictive uncertainty in Laplace approximations depends critically on the prior precision, yet selecting it requires a posterior sweep that is prohibitively expensive for neural networks with billions of parameters. Under the Maximal Update Parametrization (μP), we derive a rescaling of the prior covariance that makes the selected precision stable as model width grows. This leads to σTransfer: we select the precision on a smaller model and zero-shot transfer it to the much larger model, i.e., without searching for the precision on the larger model at all. We show convergence of the prior kernel, posterior covariance, selected precision, and posterior-derived decisions under explicit conditions, and verify σTransfer across regression, image classification, and Transformer readouts. For example, measured precision-sweep speedups reach \sim 5000\times when transferring from width 128 to 4096 on MNIST, at a target-NLL degradation of 0.002; transferring from a public 1B to 7B model gives a median search speedup of \sim 2.3\times (up to \sim 330\times), with a mean measured target-NLL increase below 10^{-4} across ten tasks. The same posterior stability also enables transfer of acquisition, OOD-detection, and abstention decisions without constructing a target posterior.