Precision is vital for robotic manipulation. Language reasoning (VLAs) are limited in their precision. Pixel generation (WAMs) are overkill.
GAMs are an alternative to both. It uses 3D grounding models instead of language models of video generators.
VLAs? WAMs? Are language and video models the right foundations for robotics?
Introducing Grounded Action Model (GAM): a new paradigm that builds robot foundation models on top of a pretrained 3D grounding model.
Ground first. Then learn to act. 🧵👇
00:00



