It turns out multi step backpropaganda is better.
paper has a beautiful way of improving backpropagation. One iteration cleanly gets us backprop, multiple iterations get us a preconditioned update.
We built high-throughput materials labs in Menlo Park to create a loop between experiments and models. The labs generate fresh data, the models learn from it, and then help us decide what to try next.
Using only 1,300 H200s, plus months of our experimental data, we mid-trained
Some news: 5 months ago I joined @coreauto, a small research lab focused on new ways for models to learn. I’m also hiring four interns to do something my younger self would've loved.
It started with a pretty unusual project. We set up a handful of small online businesses, and