Pinned
One question that's been on my mind for years now is: could we use regular multimodal LLMs not necessarily trained for robotics to do the high level robotics intelligence part that VLAs and WAMs attempt to do?
The latest explosion of powerful opensource multi-modal LLMs has,




