I decided to review and explicitly post about the limitations of all my previous papers because I believe this is the fundamental driving force behind research, especially in this era of PRs and bubbles.
We’re building a foundation motor control policy 🧠 that can control robots on the ground 🦿and in the air 🚁.
It is trained on millions of embodiment variations 🤖, and we are seeing exciting transfer to unseen robots.
We’ve collected 200+ real-world robot models to shape the
Happy to share what folks have been building! Kudos to Nico @NicoBohlinger and Bo @BoAi0110 who led this research on multi-embodiment generalist learning!
⚡ One policy, millions of embodiments, over 200 robot models. Can we add yours?
We're building γ₀, a generalist RL policy for motion control trained across millions of randomized embodiments derived from a growing collection of more than 200 robot models.
Automatic curriculum became fun in my early research on general human motion tracking in FLD arxiv.org/abs/2402.13820, where one has to learn from a continuously parameterized (infinite) motion dataset.
In this case, it became crucial to determine which motions to learn first.
We achieved high-speed rough-terrain locomotion on ANYmal with an automatic curriculum.
LP-ACRL automatically samples terrain types, levels, and velocity commands at the correct time based on policy performance, without predefining an order.
🔗sites.google.com/view/lp-acrl
🧵RAL 2026
This is actually very related! It is actually a bandit problem for the task scheduler (teacher) based on the signal (learning progress) it receives.
An extension to the continuous-time bandit in human motion learning is exemplified in FLD (arxiv.org/pdf/2402.13820) Section A.2.6!