Log inSign up
Wenlong Huang
741 posts
Wenlong Huang profile banner
@wenlong_huang

Wenlong Huang

@wenlong_huang
PhD Student @StanfordSVL. Previously @MIT_CSAIL @Berkeley_AI @GoogleDeepMind @NVIDIARobotics. Robotics, Foundation Models, Spatial Intelligence.
Stanford, CA
wenlonghuang.com
Joined May 2019
1,421
Following
5,887
Followers
RepliesRepliesRepostsRepostsMediaMedia

Log in or sign up for X

See what’s happening and join the conversation

Continue with phone
or
Log in with username or email
Terms·Privacy·Cookies·Accessibility·Ads Info·© 2026 X Corp.
  • Pinned
    @wenlong_huang
    Wenlong Huang
    @wenlong_huang
    Jan 8
    What if we can simulate an *interactive 3D world*, from a single image, in the wild, in real time? Introducing PointWorld-1B: a large pre-trained 3D world model that predicts env dynamics given RGB-D capture and robot actions. 🌐 point-world.github.io from @Stanford @nvidia
    Image
    00:00
    31
  • @wenlong_huang
    Wenlong Huang
    @wenlong_huang
    Aug 26
    @deepakpathak always has the best vision on how to make robots generalize, which I learned tremendously from during my years working with him (and continuously so!). And video prompting does feel like a key unlock towards in-the-wild robots. Huge congratulations to the team
    @SkildAI
    Skild AI
    @SkildAI
    Aug 25
    Introducing S1, our new foundation model that learns from one example. It can be taught 10-minute long tasks that it has never seen before, from one video prompt without any fine-tuning. Watch S1 operate in real-time via in-context learning:
    Image
    00:00
    3
  • @wenlong_huang
    Wenlong Huang
    @wenlong_huang
    Jul 24
    What should be the actions when robots are learning from videos? Our answer: use 𝗮𝗰𝘁𝗶𝗼𝗻𝘀 𝗶𝗻 𝘁𝗵𝗲 𝗹𝗮𝗻𝗴𝘂𝗮𝗴𝗲 𝗼𝗳 𝘃𝗶𝘀𝗶𝗼𝗻. Introducing 𝗠𝗮𝘀𝗸𝗲𝗱 𝗩𝗶𝘀𝘂𝗮𝗹 𝗔𝗰𝘁𝗶𝗼𝗻𝘀 (𝗠𝗩𝗔), a technique to turn pre-trained video models into a robot world model
    @HadiZayer
    Hadi Alzayer
    @HadiZayer
    Jul 24
    We trained a video world model on just 15 hours of video of a single-arm robot. 🧵 It generalizes zero-shot to unseen embodiments (and even orangutans). And it picked up something we never trained for: give it the object motion you want, and it synthesizes the robot motion that
    Image
    00:00
    1
  • @wenlong_huang
    Wenlong Huang
    @wenlong_huang
    Jul 21
    Very exciting news! Congratulations @YunzhuLiYZ and @drfeifei !!
    @drfeifei
    Fei-Fei Li
    World Labs
    @drfeifei
    Jul 21
    The world is not just made of words, and spatial intelligence was never just about perceiving and generating worlds. It's about interacting with them. Today, SceniX is joining World Labs. 🌎🤖👇
    Image
    00:00
  • @wenlong_huang
    Wenlong Huang
    @wenlong_huang
    Jul 16
    Though not attending RSS in-person this year, I will give a remote talk and join the panel discussion at the Data-Centric Robotics Workshop today. I will talk about “Scaling Robotics with Counterfactuals”, a different take on how robot data should be scaled, and one that might
    1
Advertisement
Advertisement