- This is going to be a super fun workshop on scaling & learning dynamics & optimization! Please consider submitting your best work there!The High-Dimensional Learning Dynamics Workshop @ ICML 2026 🇰🇷, with a special focus on scaling laws, is coming up, July 10! Submission Deadline: May 11 AoE (extended). @Locchiu @albertobietti @JustinLin610 @inbarser @BachFrancis @ShamKakade6 @andrewgwils @blake__bordelon
- The originality and the depth of science are really impressive. High thinking, signal to flop ratio. Congrats Damien, Elliot, Courtney et al.1/10 We built ADANA, an optimizer that gets better as you scale. It extends AdamW with log-time schedules for momentum and weight decay — same hyperparameter count, no extra engineering. Scaled from 45M to 2.6B, it saves ~40% compute vs tuned AdamW, and the gap keeps growing.🧵
- First scaling law: performance follows a power law of compute, with its exponent governed by science and engineering. Second scaling law: the total improvement of this law follows a power law of resource, with its exponent governed by vision and conviction.
- come and enjoy the blessing from universalityRealistic training dynamics are too complex to be described by simple scaling laws with hand-picked formulas, yet they obey precise scaling trends and universality. Join me Tue morning at the Theory and Phenomenology oral session for an alternative approach that gets us farther!






