Deep neural networks exhibit regular macroscopic behavior despite highly nonlinear dynamics in vast parameter spaces. We develop a statistical-mechanical description of learning directly in function space, treating parameter configurations as microscopic realizations and functions with their dynamical operators as macroscopic variables. For mean-squared loss, the exact error dynamics are governed by the learning operator M=JJ^\ast. Combining the dynamical Boltzmann weight of the conditional stochastic dynamics with the parameter-space density of states, whose local curvature defines a statistical operator B, and integrating over local fluctuations yields $ Φ_{fluc}(M;B)=\frac{σ_ξ^2}{2}\log\det(M^{-1}+B)+const. $ At fixed spectrum, this term is rotationally stationary when [M,B]=0, is minimized by pairing large eigenvalues of M with small eigenvalues of B, and generates a local restoring contribution against rotational mismatch. For ReLU-type function spaces under mild stable statistical conditions, B=σ_ξ^2L^\ast\mathcal K L, where L measures coarse-grained second-order structure. Thus the low-B sector corresponds, up to bounded anisotropy of \mathcal K, to low structural curvature, implying a preference for faster relaxation along smooth, data-adaptive directions. These results identify function space as a natural macroscopic level for studying stable collective organization in learning.
A Function-Space Approach to the Statistical Mechanics of Learning Dynamics
Deep neural networks exhibit regular macroscopic behavior despite highly nonlinear dynamics in vast parameter spaces. We develop a statistical-mechanical description of learning directly in function space, treating parameter configurations as microscopic realizations and…
- Preview

- Year
- 2026
- Hosting
- Abstract onlyARXIV-DEFAULT
Cite
Notes
Only stored in your browser.
Attribution
- Abstract & full text
- arxiv.org/abs/2609.09589ARXIV-DEFAULT
- TL;DR
- Semantic Scholar