Pinned
New research from our MLO Lab @EPFL: Improving Neural Network Training by Decoupling the Magnitude and Direction of Weight Vectors.
Magnitude-Direction Decoupling (MD): a simple optimizer tweak, and what we (currently) think is the right way to train efficiently at scale. 🧵


