Pinned
Recursive Self-Improvement through Multi-agent RL Post-training and Unsupervised Environment Design (UED)… but it actually works!
Delighted to finally release this paper, which trains a single LLM to act as both an Environment Designer to build new multi-turn RL training
Continuous self-improvement needs an ever-expanding supply of training environments (goals).
SPADE: one model self-plays the Environment Designer and the Reasoning Agent, writing executable, agentic environments that get harder as it improves. Environment scaling on its own. ♠️
00:00





