Volkan Cevher, Thomas Michaelsen Pethick, Wanyun Xie
Sharpness-aware minimization (SAM) has been shown to improve the generalization of neural networks. However, each SAM update requires sequentially computing two gradients, effectively doubling the per-iteration cost compared to base optimizers like SGD. We ...
2024