I gave a talk at ICLR 2026 about how we are scaling RL on frontier LLMs with 1T+ parameters, on experimental data from our physical lab at Periodic!
Here's a rough recording of the talk:
Good blog, makes you think about the empirical observation that cureent RL methods that work for LLMs are *low bias*
- value functions trade off variance with bias, and hasn't shown huge gains yet
- small bias from trainer-inference mismatch often is catastrophic for scaled up