Current video models have no reliable way to guarantee correct physics in the generated output. Other than infinitely scaling up video models and hoping physical properties “emerge”, can we do something different?
Here is our recent exploration: we bridge physical simulators and
The key idea is that physics simulator provides coarse prediction in 3D space and renders it into optical flows and RGB previews, and a video generator consumes both signals to “render” it into a realistic video output.
Okay, now FID50k is almost 0.7 with a bigger model and I think 0.6 or 0.5 is just a matter of time. But our point is that FID MEANS NOTHING AND YOU SHOULD NOT CHASE IT. Not sure if this worth a paper but will make it public soon!
Check out our new progress on SfM: Instantsfm now makes large-scale (thousands of images) way more efficient 🚀 Huge engineering effort from our amazing teammates!
🚀 Introducing InstantSfM: Fully Sparse and Parallel Structure-from-Motion.
✅ Python + GPU-optimized implementation, no C++ anymore!
✅ 40× faster than COLMAP with 5K images on single GPU!
✅ Scales beyond 100 images (more than VGGT/VGGSfM can consume)!
✅ Support metric scale.
Thanks to all my collaborators, and shoutout to Yue! So grateful for such an incredibly supportive team – it’s truly a pleasure working with all of you!