Scaling Regimes for Two-Layer Neural Networks

Course project · MATH 562, Theory of Machine Learning, McGill University

A two-layer network's behavior as width grows depends on a single scaling factor in front of its output. I compared the three standard choices, Neural Tangent Kernel, Mean Field, and Random Features, across three datasets (synthetic, MNIST, CIFAR-10) and a range of widths, activations, and learning rates.

  • Mean Field reached the lowest test loss across the board, consistent with its learning rate scaling with width.
  • Random Features, which only trains the outer weights, performed worst, as expected from learning fewer features.
  • NTK landed in between, with output variance that stays stable rather than shrinking to zero as width grows.

Analysis of Neurological Signals for Exoskeleton Control

Course project · WCOM 206, McGill University

A writing-course paper comparing three signals used to detect movement intent for exoskeleton control in neurorehabilitation: EEG, surface EMG, and a fusion of the two. I compared accuracy, speed, and equipment cost across studies using each one.

  • sEMG won on all three (96% accuracy, 59.9ms, ~US$672 in sensors), since most of EEG's cost comes from cleaning up its noise.
  • EEG + sEMG fusion is slower and pricier, but more robust early in rehabilitation when a patient's sEMG signal is still weak.
  • The studies compared aren't directly comparable (different subject counts, classes, equipment), so the conclusions are more suggestive than definitive.