Lenka Zdeborová, Florent Gérard Krzakala, Bruno Loureiro, Hugo Chao Cui, Yatin Dandi, Luca Pesce, Yue Lu
In this manuscript, we investigate the problem of how two-layer neural networks learn features from data, and improve over the kernel regime, after being trained with a single gradient descent step. Leveraging the insight from (Ba et al., 2022), we model t ...
2024