Did Anthropic get more gains out of model scaling than other labs thought was possible? It reminds me of an interesting recent paper, which showed that deep layers in open LLMs are not doing much, and that this can be fixed by scaling the LayerNorm output.
The recent Microsoft AI report noted that too much learning rate decay during pretraining hurts post-RL performance. This is actually just the latest of several papers this year pointing out that small learning rates can be harmful in LLM pretraining. (Thread)
Looking forward to attending ICLR and giving a talk on Sunday at 9am at the Science of Deep Learning workshop: scienceofdlworkshop.github.io/2026/. Message me if you want to chat about deep learning optimizer dynamics at the conference!
AMI Labs founder Yann LeCun on why LLMs are fooling us the same way AI has for decades:
He argues that every generation of AI scientists has made the same mistake: confusing task performance with real intelligence.
LeCun's core challenge to the current hype:
"We're fooled into