Do RL solutions share a common structure? We show that all solutions of Reinforcement Learning lie on a hyperplane. Our work, Proto Successor Measure, learns this abstraction of the MDP to do zero-shot RL for any reward function. (1/n)
Reminder! RLBRew deadline in coming up in 7 days! Submit your works soon👩💻
Reminder that we accept under review papers! This is a good place to discuss your ideas and get feedback from the community
Introducing RLDP – a simple, scalable approach for building strong Behavioral Foundation Models. 🚀 #ICLR2026
✅ Robust objective: avoids brittle unsupervised RL objectives while staying simple and scalable.
🗂️ No data-coverage limitations: works across a wide variety of
I’ll be attending #NeurIPS2025 and presenting our work, “RLZero: Direct Policy Inference from Language Without In-Domain Supervision." Excited to chat about unsupervised RL, reasoning, and RL more broadly. I’m also exploring industry opportunities — feel free to reach out!
I’ll be at #ICML2025 presenting our paper, “Proto Successor Measure: Representing the Behavior Space of an RL Agent”. Excited to connect with others working on unsupervised RL and RL more broadly. Also am on the lookout for research collaborations and opportunities in industries.