1. X
  2. Yunhao (Andy) Ge
Log inSign up
Yunhao (Andy) Ge
143 posts
Yunhao (Andy) Ge profile banner
user avatar

Yunhao (Andy) Ge

@GeYunhao
Research Scientist @NVIDIA GEAR Lab | CS PhD @USC, Ex Visiting PhD @Stanford, Amazon ML Fellow @Amzaon, intern @Google, @Microsoft | VLA, World Foundation Model
Santa Clara
gyhandy.github.io
Joined January 2021
212
Following
535
Followers
RepliesRepliesMediaMedia
  • Pinned
    user avatar
    Yunhao (Andy) Ge
    @GeYunhao
    Feb 4
    Words in. Worlds imagined. Actions out. 🤖🌎 DreamZero lets robots dream in pixels and act—via joint video + action prediction. 🔥2× better generalization than VLAs ⚡14B @ 7 Hz 🤝Cross-embodiment transfer (w/ 10–20 min video) 🦾New robot, 30 min play, zero-shot skills intact
    user avatar
    Joel Jang
    @jang_yoel
    Feb 4
    Introducing DreamZero 🤖🌎 from @nvidia > A 14B “World Action Model” that achieves zero-shot generalization to unseen tasks & few-shot adaptation to new robots > The key? Jointly predicting video & actions in the same diffusion forward pass Project Page: dreamzero0.github.io
    Image
    00:00
  • user avatar
    Yunhao (Andy) Ge
    @GeYunhao
    Jul 15
    RoboTTT🤖 uses Test-Time Training with fast weights to tackle long-horizon robotic tasks, scaling visuomotor context to 8K timesteps—1,000× longer than prior policies—without increasing inference latency. Amazing work!🚀
    user avatar
    Yunfan Jiang
    @YunfanJiang
    Jul 15
    We scaled robot policies to 8K timesteps of visuomotor context, orders of magnitude beyond current SoTAs, at constant inference latency. Introducing RoboTTT 🤖 With minutes of experience in context, our robots: 🎥 one-shot imitate human video demos 📈 improve themselves during
    Image
    00:00
  • user avatar
    Yunhao (Andy) Ge
    @GeYunhao
    Mar 23
    Honored to contribute to DreamZero & Cosmos Policy 🤖 Seeing them featured in Jensen’s keynote made it all feel real — super proud of the team. GR00T N2, let’s GOOOO! 🚀
    Image
    Image
  • user avatar
    Yunhao (Andy) Ge
    @GeYunhao
    Feb 25
    Human video is the most scalable source of physical intelligence. EgoScale answers a fundamental question: human video exhibits a clear scaling law for dexterous manipulation.
    user avatar
    Jim Fan
    @DrJimFan
    Feb 25
    We trained a humanoid with 22-DoF dexterous hands to assemble model cars, operate syringes, sort poker cards, fold/roll shirts, all learned primarily from 20,000+ hours of egocentric human video with no robot in the loop. Humans are the most scalable embodiment on the planet. We
    Image
    00:00
  • user avatar
    Yunhao (Andy) Ge
    @GeYunhao
    Apr 24, 2025
    DreamDistribution: Learning Prompt Distribution for Diverse In-distribution Generation 🎤 Catch @briannlongzhao at #ICLR2025 Poster Session 1 (#177) on Apr 24! 📄 Paper + Code + Demo: briannlongzhao.github.io/DreamDistribut…
    user avatar
    Brian Zhao
    @briannlongzhao
    Apr 23, 2025
    Introducing DreamDistribution, a novel, simple approach to learn semantic distribution over multiple input images, enabling generation of diverse in-distribution images and 3D with editing flexibilities. We will be presenting our work at @iclr_conf Apr 24 poster session 1 #177
    Image
    00:00

Log in or sign up for X

See what’s happening and join the conversation

Continue with phone
or
Log in with username or email
Terms·Privacy·Cookies·Accessibility·Ads Info·© 2026 X Corp.
Advertisement
Advertisement