To bring generalist intelligent robots to the real world, we have to overcome the data scarcity problem.
At Rhoda, we are solving it by reformulating robot policies as video generation.
Today, we introduce the Direct Video-Action Model (DVA)
At Rhoda, we care deeply about the science of pre-training for robotics.
In one of the most rigorous studies of its kind, over thousands of trials and hundreds of hours of robot evaluations, we show how scaling web-video pre-training leads to better real-world robot performance.
Does scaling pre-training on general web video improve a complex manipulation task in real deployment?
We scale model size and pre-training compute, and test on one industrial task.
Yes. The better a pre-trained model predicts web video, the better its post-trained policy. 🧵
Can a large foundation video model run as a real-time robot policy at the edge, on a single RTX 5090?
• ✅ No quantization
• ✅ No distillation
• ✅ Full denoising (all the way from noise to clean video)
We just proved it's possible. 👇🎬
Teaching a robot a new task typically means stopping operations, collecting teleoperated demonstrations, and retraining. That process takes hours at a minimum. We wanted to know if we could collapse it to seconds — from a single human demo, on the fly, no retraining required.