A primary advantage of neural networks lies in their feature learning characteristics, which is challenging to theoretically analyze due to the complexity of their training dynamics. We examine feature learning and its potential benefits for generalization from a statistical perspective. After reviewing the neural tangent kernel (NTK) theory and recent results in kernel regression, which address the generalization issue of sufficiently wide neural networks, we examine limitations and implications of the fixed kernel theory (as the NTK theory) and review recent theoretical advancements in feature learning. Moving beyond theories with fixed features, we consider neural networks as adaptive feature models. Finally, we propose an over-parameterized Gaussian sequence model as a prototype for the adaptive feature model to study feature learning characteristics and motivate their future analysis for neural networks.
Towards a Statistical Understanding of Neural Networks: Beyond the Neural Tangent Kernel Theories
A primary advantage of neural networks lies in their feature learning characteristics, which is challenging to theoretically analyze due to the complexity of their training dynamics.
- Year
- 2024
- Hosting
- Abstract onlyARXIV-DEFAULT
Cite
Notes
Only stored in your browser.
Attribution
- Abstract & full text
- arxiv.org/abs/2412.18756ARXIV-DEFAULT
- TL;DR
- Semantic Scholar