If you are interested in TTT (and its limitations) at all, you should talk to Junchen!
HALL A #2910
Thursday 10:30 AM – 12:15 PM
- We are presenting today in the afternoon poster session (15:30-17:30) at Poster No. 28! #CVPR2026🚀 Exciting news! We’re introducing VGG-T³: a scalable model for offline feed-forward 3D reconstruction that finally tackles the "quadratic bottleneck." Ever wanted to have VGGT reconstruct a 1,000-image scene in seconds instead of 10 minutes and use it for visual localization?
- Traditional 3D reconstruction pipelines like COLMAP operate in a loop, growing the scene piece-by-piece. 🧩🔄 With DejaView, we introduce this inductive bias for feed-forward 3D reconstruction—running a single alternating-attention block in a loop! Awesome work from the team 👇Do 3D reconstruction transformers really need a billion parameters, or are most of those layers just doing the same thing over and over? Introducing Déjà View: a single transformer block, looped K times, that matches or beats models 8–10× its size with lower compute. 🧵
- We just released code and model! Go check it out! Code: github.com/nv-dvl/vgg-ttt Model: huggingface.co/nvidia/vgg-ttt🚀 Exciting news! We’re introducing VGG-T³: a scalable model for offline feed-forward 3D reconstruction that finally tackles the "quadratic bottleneck." Ever wanted to have VGGT reconstruct a 1,000-image scene in seconds instead of 10 minutes and use it for visual localization?
- 🚀 Exciting news! We’re introducing VGG-T³: a scalable model for offline feed-forward 3D reconstruction that finally tackles the "quadratic bottleneck." Ever wanted to have VGGT reconstruct a 1,000-image scene in seconds instead of 10 minutes and use it for visual localization?



