our new system trains humanoid robots using data from cell phone videos, enabling skills such as climbing stairs and sitting on chairs in a single policy
(w/ @redstone_hong@junyi42@davidrmcall)
On overfitting being helpful, I don’t think it’s unique to robotic BC
eg. from InstructGPT “we find that our SFT models overfit on validation loss after 1 epoch; however, we find that training for more epochs helps both the RM score and human preference ratings”.
I think the
Behavioral cloning mystery
seohong.me/blog/behaviora…
I wrote a new blog post about "mysteries" in behavioral cloning that appear with real-world robot data (e.g., overfitting is "good"). I also tried to demystify them and shared my thoughts!
Hand teleop is an import yet challenging problem, Adam / Sarthak's project is a really great step towards making such systems accessible and reliable, check it out!
Excited to share SPD: simulation pre-training for dexterity. We pre-trained a policy in simulation and fine-tuned with less than 2 hours of real data (with @sarthakkamat)
Introducing ABC: open data, training, and infrastructure for robotics.
We release the largest teleop dataset to date, and extensively investigate design decisions, pretraining, and post-training techniques.
@arthurallshire@Cinnabar233@adamrasb@redstone_hong@davidrmcall
if you take GPU maxis at their word, they’re the swiss army knife of computing
but if they can be easily slotted into any training job, their deployments have no stickiness bc a competitor can also easily replace them
they have to monopoly win or margins go to zero
bearish...
if you take humanoid maxis at their word, they’re the swiss army knife of robots
but if they can be easily slotted into any role, their deployments have no stickiness bc a competitor can also easily replace them
they have to monopoly win or margins go to zero
bearish