1. X
  2. Min-Hung (Steve) Chen
Log inSign up
Min-Hung (Steve) Chen
1,034 posts
Image
user avatar
Min-Hung (Steve) Chen
@CMHungSteven
Staff Research Scientist, NVR TW @NVIDIAAI @NVIDIA (Project Lead: DoRA, EoRA, 4D-RGPT, SpatialClaw) | Ph.D. @GeorgiaTech | github.com/cmhungsteve
Taipei City, Taiwan
minhungchen.netlify.app
Joined July 2011
1,696
Following
2,524
Followers
RepliesRepliesMediaMedia
  • Pinned
    user avatar
    Min-Hung (Steve) Chen
    @CMHungSteven
    May 11, 2022
    (1/N) Are you looking for #Vision #Transformer papers in various areas? Check out this list of papers including a broad range of different tasks! github.com/cmhungsteve/Aw… Feel free to share with others😀 @Montreal_AI @machinelearnflx @hardmaru @ak92501 @arankomatsuzaki @omarsar0
    Image
    GitHub - cmhungsteve/Awesome-Transformer-Attention: An ultimately comprehensive paper list of...
    From github.com
  • user avatar
    Min-Hung (Steve) Chen
    @CMHungSteven
    Jul 6
    2 denoising steps can capture better physics than the full 50-step video. 🤯 That's the insight behind PhaseLock: it locks that early motion prior back into the final high-fidelity output via Latent Δ Guidance — training-free, +6.2 pts physical consistency, ~zero extra compute
    Image
    00:00
    Image
    user avatar
    Woojung Han
    @woojung0305
    Jul 6
    ᴅᴏ ᴠɪᴅᴇᴏ ᴅɪꜰꜰᴜꜱɪᴏɴ ᴍᴏᴅᴇʟꜱ ʀᴇᴀʟʟʏ ɴᴏᴛ ᴋɴᴏᴡ ᴘʜʏꜱɪᴄꜱ, ᴏʀ ᴅᴏ ᴛʜᴇʏ ꜰᴏʀɢᴇᴛ ɪᴛ ᴅᴜʀɪɴɢ ɢᴇɴᴇʀᴀᴛɪᴏɴ? 🧐🎬 Want to know more? Please drop by my #icml poster presentation at tomorrow's morning session (Hall A #905!)
  • user avatar
    Min-Hung (Steve) Chen
    @CMHungSteven
    Jun 23
    Author here — huge thanks to @NVIDIAAI for sharing SpatialClaw 🙏 What excites me most: code becomes the agent’s workspace 🐍 SpatialClaw lets a VLM compose perception tools, inspect intermediate evidence, and revise before answering. #NVIDIAAI #SpatialAI
    Image
    00:00
    Image
    user avatar
    NVIDIA AI
    NVIDIA
    @NVIDIAAI
    Jun 16
    Code is the right action interface for spatial reasoning agents. New from NVIDIA Research: SpatialClaw, a training-free agent that uses code as its action interface for complex visual tasks. Instead of calling a fixed set of pre-defined tools, the agent writes Python inside a
  • user avatar
    Min-Hung (Steve) Chen
    @CMHungSteven
    Jun 3
    T4V @CVPR is starting!!
    Image
    Image
    Image
    user avatar
    Min-Hung (Steve) Chen
    @CMHungSteven
    May 25
    What comes after today’s visual backbones? At T4V @CVPR 2026, we’re bringing the community together for a focused half-day workshop on Transformers for Vision and Multimodal AI — covering image, video, 3D, MLLMs, efficient attention, SSMs/Mamba, and the next generation of visual
  • user avatar
    Min-Hung (Steve) Chen
    @CMHungSteven
    Jun 3
    T4V @CVPR starts today🚀🚀 1:45–5:40pm, Room 607 Join talks/discussion on Transformers for image, video, 3D, MLLMs, efficient attention & SSMs/Mamba
    user avatar
    Min-Hung (Steve) Chen
    @CMHungSteven
    May 25
    What comes after today’s visual backbones? At T4V @CVPR 2026, we’re bringing the community together for a focused half-day workshop on Transformers for Vision and Multimodal AI — covering image, video, 3D, MLLMs, efficient attention, SSMs/Mamba, and the next generation of visual
    Image

Log in or sign up for X

See what’s happening and join the conversation

Continue with phone
or
Log in with username or email
Terms·Privacy·Cookies·Accessibility·Ads Info·© 2026 X Corp.
Advertisement
Advertisement