Glad to share that our work, VGGT-Det (CVPR 2026). Existing multi-view indoor 3D detectors rely on precisely calibrated multi-view camera poses or depth—costly to obtain in real world. VGGT-Det changes this with a Sensor-Geometry-Free framework (no multi-view poses or depth)