Log inSign up
Vincent Sitzmann
959 posts
@vincesitzmann

Vincent Sitzmann

@vincesitzmann
Building AI that learns by interacting with the world. Associate Professor @ MIT, leading the Scene Representation Group (scenerepresentations.org).
Cambridge, Massachusetts
vincentsitzmann.com
Joined February 2016
329
Following
19.9K
Followers
RepliesRepliesRepostsRepostsMediaMedia

Log in or sign up for X

See what’s happening and join the conversation

Continue with phone
or
Log in with username or email
Terms·Privacy·Cookies·Accessibility·Ads Info·© 2026 X Corp.
  • Pinned
    @vincesitzmann
    Vincent Sitzmann
    @vincesitzmann
    Jun 8
    Introducing MilliVid, our new method for long-context video generation! MilliVid creates videos that are consistent over long time spans, without using retrieval heuristics or 3D maps! (1/n) davidcharatan.com/millivid/#
    Image
    00:00
    11
  • @vincesitzmann
    Vincent Sitzmann
    @vincesitzmann
    Sep 8
    At 10:20 am Malmö time, I will be speaking at the "X-Reason" workshop in the Palisades South room (turn right at registration, then down the stairs) about whether intermediate representations are important for embodied intelligence! #ECCV2026
    2
  • @vincesitzmann
    Vincent Sitzmann
    @vincesitzmann
    Sep 7
    I agree with Phil's take here: The progress of LLMs on controlling robots is quite interesting. Intuitively, this makes sense: controlling a robot is not so different from computer use, and an agent that is good at computer use is probably also good at controlling a robot and
    @phillip_isola
    Phillip Isola
    @phillip_isola
    Sep 7
    Recently, there have been a lot of impressive demos of AI agents, like Claude, controlling robots. I wrote a short blog post with my thoughts on the advent of these "robot-use agents." web.mit.edu/phillipi/www/w… I think it's an important change in the trajectory of robotics!
    8
  • @vincesitzmann
    Vincent Sitzmann
    @vincesitzmann
    Sep 3
    This looks like a really cool product, congrats to the @theworldlabs team!
    @theworldlabs
    World Labs
    @theworldlabs
    Sep 1
    Introducing Atlas: The world's first multimodal world model that generates image and video frames with pixel-perfect camera control and reconstructs them in 3D. Model the world, move the camera, and simulate space & time.
    Image
    00:00
    5
  • @vincesitzmann
    Vincent Sitzmann
    @vincesitzmann
    Aug 25
    I was a guest on the @MIT_CSAIL podcast to discuss robots, AI, and what the near future might look like. @klgiven did a great job steering the conversation, and I think much of our chat is accessible to non-experts - hope you find it interesting! bit.ly/4zruZdt Also on
    Image
    Why AI Still Can't Load the Dishwasher | CSAIL Alliances
    From cap.csail.mit.edu
    1
Advertisement
Advertisement