You're absolutely right, I should not have allowed that mineshaft to collapse with workers still inside, and that's on me. They were the lode-baring component—
This is great work! And I finally get to live my dream of Schmidhubering, just once:
Using the average of rollouts from the same state as a baseline is exactly what I did in the original DDPO paper (which predates GRPO). But I thought it was boring so I put it in the Appendix :)
Interaction with the real world is the major bottleneck in robot learning. So what would robot RL look like if we didn’t need to limit compute per interaction? Our latest work, Off-Policy Generative Policy Optimization (OGPO, accepted to ICML26) embarks on answering this question
π0.7 handles diverse prompts that don't just say what to do, but also how to do it, including rich language and multimodal information, such as visual subgoal images. At test time, these images can be produced by a lightweight world model.
oh you're using VLAs? everyone's using GRPs now. just kidding we're all on LBMs. world models are the future so we developed our own WAM. we're using DVAs. we were using UWMs but our robot caught on fire so we switched to DreamUMVLAPs. we're shipping a robot that passes butter.