Cooperative language-grounded scene understanding over independently reconstructed Gaussian maps.
🌟 Overview · 🎬 Demo · 🧩 Framework · 🚀 Release Status
🚧 Code · Dataset · Pretrained Models — Coming Soon
CoRef-GS enables multiple robots to independently construct local semantic Gaussian maps and integrate them into a shared representation for cooperative referring scene understanding.
It preserves language-grounded referring ability across independently reconstructed and aligned maps, while interpreting spatial relations from the querying robot's viewpoint.
Overview of CoRef-GS for cooperative multi-agent referring scene understanding.
A visual overview of the CoQuad-Ref dataset, multi-agent semantic Gaussian mapping, cross-agent map alignment and fusion, and cooperative referring results.
The demo showcases selected components of CoRef-GS, including our real-world and simulated data, the overall model framework, cooperative map construction and fusion, and qualitative referring results across different robot viewpoints.
Overall framework of CoRef-GS.
🚧 Code, dataset, and pretrained models will be made publicly available.
- 🎬 Release the CoRef-GS demo video.
- 💻 Release the source code.
- 📚 Release the CoQuad-Ref dataset.
- 🧠 Release the pretrained models.

