That new LFM2.5-350M is super overtrained, right? And everyone was shocked about how far they pushed it?
As it turns out, we have a brand new scaling law for that! 🧵
[1/n]
A 350M model overtrained to 80,000+ tokens per parameter should be a mistake. Nicholas Roberts' ( @nick11roberts ) scaling law says it's optimal.
The Train-to-Test (T²) law folds pass@k into pretraining scaling, so once a model reasons at test time, being compute-optimal flips
Great read and congrats on GLM-5.3! Here’s my paper that @jietang cites—optimal tokens/param is task-dependent: memorization wants params, reasoning wants data arxiv.org/abs/2503.10061
He also mentions inference/reasoning cost+scaling. We do this with T² arxiv.org/abs/2604.01411
Thoughts About Scaling Law
Scaling, but not only of parameters. Every model release now ends with the same question: how many parameters? It isn't a question that can be answered on its own. Parameter count is only meaningful alongside three others — how much data you have,
Two Reading Groups, one week 👀
Monday, August 17 at 4:00 p.m. PT: @nick11roberts (Incoming Postdoctoral Research Fellow, @Princeton) will present Train-to-Test (T²) Scaling Laws, exploring how inference costs can radically shift optimal pretraining into the overtraining regime.