Pinned
Scaling laws have transformed how we understand pretraining, and recent work has begun to characterize scaling in RL. But these stages are almost always studied in isolation. Can we study pretraining and RL jointly, and derive a unified scaling law for the full pipeline? 🧵



