1. X
  2. Chenming Zhu
Log inSign up
Chenming Zhu
13 posts
user avatar
Chenming Zhu
@chenming_eric
Ph.D. Candidate @ HKU-MMLab | 3D Vision & Language for Robotics πŸ€–
Hong Kong
zcmax.github.io
Joined December 2019
111
Following
18
Followers
RepliesRepliesMediaMedia
  • user avatar
    Chenming Zhu
    @chenming_eric
    Dec 18, 2025
    MMSI-Video-Bench is the most challenging video spatial intelligence benchmark, which can also be used to test the world model, such as Veo3. πŸ˜ŠπŸ‘ Also, we introduce our online video spatial intelligence benchmark in streaming settings for #NeurIPS2025: OST-Bench πŸš€
    user avatar
    Runsen Xu
    @runsen_xu
    Dec 16, 2025
    1/3 Introducing MMSI-Video-Bench, the most challenging Spatial Intelligence benchmark. Even the strongest MLLM, Gemini 3 Pro, scores only 38%, revealing a ~60% human–AI gap. 🀯 🌐 Project Page: rbler1234.github.io/MMSI-VIdeo-Ben… πŸ’» GitHub: github.com/InternRobotics… πŸ€— Hugging Face:
    Image
    00:00
    102
  • user avatar
    Chenming Zhu
    @chenming_eric
    Nov 28, 2025
    The most elegant framework for achieving spatial reconstruction and reasoning πŸ‘† The story and analysis is pretty worth reading
    user avatar
    Wenbo Hu
    @gordonhu608
    Nov 28, 2025
    πŸš€ Introducing G^2VLM: Geometry Grounded Vision Language Model with Unified 3D Reconstruction and Spatial Reasoning G^2VLM can natively predicts 3D attributes (depth, camera pose, pointmaps) and uses them for spatial understanding via interleaved reasoning. πŸ”§ End-to-End
    Image
    00:00
    82
  • user avatar
    Chenming Zhu
    @chenming_eric
    Aug 5, 2025
    This is crazy! I’m curious about its generalization capabilities, maybe a new approach for generating the embodied agent training data
    user avatar
    Google DeepMind
    @GoogleDeepMind
    Aug 5, 2025
    Replying to @GoogleDeepMind
    πŸ”˜ Accelerating agent research To explore the potential for agent training, we placed our SIMA agent in a Genie 3 world with a goal. The agent acts, and Genie 3 simulates a response in the world without knowing the objective. This is key for building more capable embodied
    Image
    00:00
    Image
    00:00
    Image
    00:00
    53
  • user avatar
    Chenming Zhu
    @chenming_eric
    Nov 19, 2024
    @cyodyssey ι€€εœˆεΏ«2εΉ΄οΌŒζ„Ÿθ°’Siyuanηš„ι₯­ι’±ζ‰“衏@ABCDELabs
    Image
    29
  • user avatar
    Chenming Zhu
    @chenming_eric
    Jan 6, 2024
    #CPAL2024πŸ‡­πŸ‡°
    Image
    Image
    156
  • See @chenming_eric's full profile

    Sign up
    Log in

Log in or sign up for X

See what’s happening and join the conversation

Continue with phone
or
Log in with username or email
TermsΒ·PrivacyΒ·CookiesΒ·AccessibilityΒ·Ads InfoΒ·Β© 2026 X Corp.
Advertisement
Advertisement