Matthieu Wyart, Antonio Sclocchi
Modern deep networks are trained with stochastic gradient descent (SGD) whose key hyperparameters are the number of data considered at each step or batch size B, and the step size or learning rate n. For small B and large n, SGD corresponds to a stochastic ...
National Academy of Sciences2024