New podcast/lecture combo -- a case study in the messy details of Olmo 3 post training & DPO with @scottgeng00. It's rare to make time for these discussions, but we cover:
What it takes for a research idea to make it into a (near) frontier model.
The messy side of DPO (usually
The NLP group at the University of Washington.
- Congrats to @hamishivi @yinn_oscar @RulinShao for their work on tmax 👐 credit to our data chef @yinn_oscar 🧑🍳Scaling agentic RL environments: today we're publishing 365,000+ tasks for SWE, terminal, and search agents - 23 tasksets behind one API, one sandbox lifecycle, one command.
- Automatic harness evolution appears to be a promising path toward AI self-improvement, but we find that its gains still largely come from repeated sampling and show limited generalization. Blog post: yikee.github.io/harnessevoluti… Code: github.com/rethinking-har…
- Excited to share our students are starting WAI @waiorg to do open agentic research. learn more at: wai-org.com check out their first work on open code agent recipe:
- MoEs are everywhere, but the design space is confusing: total vs active experts? expert size? shared experts? routing? token dropping? We train >2000 MoE LMs 🫠 to investigate and bring you: 📄🔪🍰 Slicing and Dicing MoEs Tl;dr: it's all about expert size and count [1/9]






