Pinned
We trained a video world model on just 15 hours of video of a single-arm robot. 🧵
It generalizes zero-shot to unseen embodiments (and even orangutans). And it picked up something we never trained for: give it the object motion you want, and it synthesizes the robot motion that



