Meet InternVLA-A1 🤖✨
It unifies scene understanding, visual foresight, and action execution into a single framework.
🧠 The Core: Synergizes MLLM's semantic understanding with world-model-style dynamic prediction, to "imagine" the future and guide adaptive actions.
🚀 The
🤖Can robots achieve accurate navigation without any external localization feedback?
📸We present #LoGoPlanner, which handles perception, localization, and planning in one go!
Check our results on LeKiWi, G1, and Go2 robots.
🌐Project: steinate.github.io/logoplanner.gi…
🚀 Introducing G^2VLM: Geometry Grounded Vision Language Model with Unified 3D Reconstruction and Spatial Reasoning
G^2VLM can natively predicts 3D attributes (depth, camera pose, pointmaps) and uses them for spatial understanding via interleaved reasoning.
🔧 End-to-End
Great release! Gallant demonstrates a clean voxel-grid pipeline for perceptive humanoid locomotion — unified policy, strong generalization across stairs, gaps, stepping stones, and cluttered spaces.
Introducing Gallant: Voxel Grid-based Humanoid Locomotion and Local-navigation across 3D Constrained Terrains 🤖
Project page: gallantloco.github.io
Arxiv: arxiv.org/abs/2511.14625
Gallant is, to our knowledge, the first system to run a single policy that handles full-space