Log inSign up
Haotong Lin
39 posts
@HaotongLin

Haotong Lin

@HaotongLin
Research Scientist at Bytedance Seed.
haotongl.github.io
Joined July 2021
283
Following
435
Followers
RepliesRepliesRepostsRepostsMediaMedia

Log in or sign up for X

See what’s happening and join the conversation

Continue with phone
or
Log in with username or email
Terms·Privacy·Cookies·Accessibility·Ads Info·© 2026 X Corp.
  • @HaotongLin
    Haotong Lin
    @HaotongLin
    Apr 24
    Congrats to the Omni team! 🥯🎉 Camera pose estimation? Just AR text gen 🤯 <campose>1.04 0.00 0.32 -0.20 -0.27 -0.02</campose> 6DoF floats as a string 🪢 no pose head, no bins, no 3D priors. On RealEstate10K: AUC@30 88.32 vs VGGT 88.23 on par with a geometry specialist!
    @CeyuanY
    Ceyuan Yang
    @CeyuanY
    Apr 24
    Introducing Omni, one unified model can support any-to-any multimodal modeling, including multimodal understanding, image/video generation and editing, world modeling and 3D reconstruction. All in one that adopts standard mixture-of-experts arch with only 3B activations.
    Image
    00:00
    2
  • @HaotongLin
    Haotong Lin
    @HaotongLin
    Apr 24
    🍌perception is a subset of generation! Strong work — excited for what comes next on the unified vision front.
    @songyoupeng
    Songyou Peng
    @songyoupeng
    Apr 23
    Yay, finally! Introducing Vision Banana🍌 from @GoogleDeepMind, our unified model that outperforms SoTA specialist models on various vision tasks! By treating 2D/3D vision tasks as image generation, we unlock a new foundation for CV. Project page: vision-banana.github.io (1/5)
    Image
    00:00
  • @HaotongLin
    Haotong Lin
    @HaotongLin
    Oct 10, 2025
    Thank you for sharing our work! Marigold is really cool! However, it’s somewhat limited by the image VAE — many flying points appear just after encoding a perfect ground-truth depth. Pixel-space diffusion to the rescue 🚀
    @AntonObukhov1
    Anton Obukhov
    @AntonObukhov1
    Oct 9, 2025
    Pixel-Perfect-Depth: the paper aims to fix Marigold's loss of sharpness induced by VAE by using VFMs (VGGT/DAv2) and a DiT-based pixel decoder to refine the predictions and achieve clean depth discontinuities. Video by authors.
    Image
    00:00
    2
  • @HaotongLin
    Haotong Lin
    @HaotongLin
    Apr 29, 2025
    Wow, thank you for crediting our work! Thrilled to see our project PromptDepthAnything being used in your latest release. This is awesome! Best of luck with the new version! !
    @ChrisAtKIRI
    Chris make some 3D scans
    @ChrisAtKIRI
    Apr 29, 2025
    Are you tired of the low quality of iPhone lidar scans? I am! And that is why we are bringing this cutting-edge iPhone lidar scan enhancement function into production! With the guidance of normal and depth, the geometry can now reach the next level! Showcases:
    Image
    00:00
    2
  • @HaotongLin
    Haotong Lin
    @HaotongLin
    Dec 19, 2024
    Check out our new work, Prompt Depth Anything, which achieves accurate metric depth estimation at up to 4K resolution! Thanks to all our collaborators!
    @bingyikang
    Bingyi Kang
    AMI Labs
    @bingyikang
    Dec 19, 2024
    Want to use Depth Anything, but need metric depth rather than relative depth? Thrilled to introduce Prompt Depth Anything, a new paradigm for accurate metric depth estimation with up to 4K resolution. 👉Key Message: Depth foundation models like DA have already internalized rich
    Image
    00:00
    2
Advertisement
Advertisement