Scaling Regimes for Two-Layer Neural Networks
Dec 2025Course project · MATH 562, Theory of Machine Learning, McGill University
A two-layer network's behavior as width grows depends on a single scaling factor in front of its output. I compared the three standard choices, Neural Tangent Kernel, Mean Field, and Random Features, across three datasets (synthetic, MNIST, CIFAR-10) and a range of widths, activations, and learning rates.
- Mean Field reached the lowest test loss across the board, consistent with its learning rate scaling with width.
- Random Features, which only trains the outer weights, performed worst, as expected from learning fewer features.
- NTK landed in between, with output variance that stays stable rather than shrinking to zero as width grows.