1. X
  2. Lerrel Pinto
Log inSign up
Lerrel Pinto
745 posts
Lerrel Pinto profile banner
user avatar

Lerrel Pinto

@LerrelPinto
Making robots more dexterous, robust and general. co-leading robotics at Meta MSL.
Menlo Park
lerrelpinto.com
Joined June 2019
217
Following
9,728
Followers
RepliesRepliesMediaMedia

Log in or sign up for X

See what’s happening and join the conversation

Continue with phone
or
Log in with username or email
Terms·Privacy·Cookies·Accessibility·Ads Info·© 2026 X Corp.
  • Pinned
    user avatar
    Lerrel Pinto
    @LerrelPinto
    Aug 20
    Strong multimodal models = strong robot models. Check out some early glimpses of Meta's Muse Spark 1.2 on robots here!
    user avatar
    Alexandr Wang
    Meta
    @alexandr_wang
    Aug 20
    1/ muse spark 1.2 is a very strong multimodal model—it can do visual coding, robotics planning, and audio-visual understanding that all come together through agentic tools.
    Image
    00:00
    Image
    00:00
  • user avatar
    Lerrel Pinto
    @LerrelPinto
    2h
    Turns out that doing In-context learning for robots is not that hard...
    Image
    00:00
  • user avatar
    Lerrel Pinto
    @LerrelPinto
    Jul 24
    Always awesome to our work be reproduced in <24 hrs ❤️
    user avatar
    atharva ☆
    @k7agar
    Jul 24
    really cool idea! reproduced patch policy on Push-T tldr is "dense ViT patch features + block causal attention > global pooled features" dense features like VLA and really high frequency. planning to test it on real hardware for some precise manipulation
    Image
    00:00
  • user avatar
    Lerrel Pinto
    @LerrelPinto
    Jul 24
    So it turns out that simply switching the visual tokens in ViTs to dense features can give us MASSIVE improvements in training efficiency, reducing model size, latency. Check the thread below for more details!
    user avatar
    Jeff Cui
    @jeffacce
    Jul 22
    Your policy doesn't need 7B params. It simply needs dense features. Introducing Patch Policy: pretrained ViT + small transformer beats OpenVLA-OFT with 0.7% of its params, and trains on a 5090. Here it inserts a cable (~2mm tol), and does it again as we unplug mid-rollout. 🧵
    Image
    00:00
  • user avatar
    Lerrel Pinto
    @LerrelPinto
    Jun 16
    Turns out you can train humanoid hands without any robot data. The idea in HUG is quite simple: (a) collect human data with smart glasses, (b) train a human manipulation model, (c) retarget to multi-fingered robot hands.
    Image
    00:00
Advertisement
Advertisement