If only the intersection area yields a viable data solution for Embodied AI, why do GraspVLA and TrackVLA — trained solely on synthetic action data — work zero-shot in the wild?
And isn’t cross-embodiment data basically a spork? Why does it outperform synthetic data as an
I wrote a fun little article about all the ways to dodge the need for real-world robot data. I think it has a cute title.
sergeylevine.substack.com/p/sporks-of-agi




