𝗔𝗳𝘁𝗲𝗿 𝟭𝟬+ 𝘆𝗲𝗮𝗿𝘀 𝗶𝗻 𝗿𝗼𝗯𝗼𝘁 𝗹𝗲𝗮𝗿𝗻𝗶𝗻𝗴, from my PhD at Imperial to Berkeley to building the Dyson Robot Learning Lab, one frustration kept hitting me:
𝗪𝗵𝘆 𝗱𝗼 𝗜 𝗵𝗮𝘃𝗲 𝘁𝗼 𝗿𝗲𝗯𝘂𝗶𝗹𝗱 𝘁𝗵𝗲 𝘀𝗮𝗺𝗲 𝗶𝗻𝗳𝗿𝗮𝘀𝘁𝗿𝘂𝗰𝘁𝘂𝗿𝗲 𝗼𝘃𝗲𝗿 𝗮𝗻𝗱
Some robot tasks are basically impossible for a human to teleoperate cleanly. Not hard, impossible. Carrying a plate across a table with two independently controlled arms, for instance. Zero percent success rate, no matter the operator.
GLIDE takes an interesting approach to
Most robots need hundreds of demonstrations to learn a task, then need hundreds more the moment anything about that task changes. This one needed one, and then just kept adapting.
The system is VLBiMan++, out of @ShenzhenUni and @DexForceAI. Show it a single human demo of
Rhoda AI checked whether scaling web video pretraining actually helps real robots, using a real customer task. Unpacking 10kg boxes of bearings and sorting the packaging.
Turns out yes. Bigger models do better, more pretraining compute does better, and the compute gap is widest
At Rhoda, we care deeply about the science of pre-training for robotics.
In one of the most rigorous studies of its kind, over thousands of trials and hundreds of hours of robot evaluations, we show how scaling web-video pre-training leads to better real-world robot performance.
A robot's camera predicts the future to decide what to do next. The problem is that prediction takes time, and time is the one thing a robot moving in the real world doesn't have to spare.
That's the tension a team from @BAAIBeijing and collaborators went after with World Action