Pinned
Check out my most recent project that I've been working on for more than half a year -- co-led with Tim @TimSong52005757!
We showed a flexible and effective recipe for improving generalist VLAs leveraging non-robotic foundation models! The key is guidance in 3D!
More details @
VLA policies learn generalist robot behaviors from massive teleoperation datasets, hoping that the right behavior emerges. But they rarely use perception during training or inference: powerful foundation models of 3D geometry, semantics, or human motion are ignored.


