Log inSign up
Sarvesh Patil
894 posts
Sarvesh Patil profile banner
@servo97

Sarvesh Patil

@servo97
Your friendly neighborhood PhD in Robotics @CMU. Dexterous Manipulation | Generative Control | Reinforcement Learning
Pittsburgh, PA
servo97.github.io
Joined November 2014
636
Following
620
Followers
RepliesRepliesRepostsRepostsMediaMedia

Log in or sign up for X

See what’s happening and join the conversation

Continue with phone
or
Log in with username or email
Terms·Privacy·Cookies·Accessibility·Ads Info·© 2026 X Corp.
  • Pinned
    @servo97
    Sarvesh Patil
    @servo97
    Jun 30
    Interaction with the real world is the major bottleneck in robot learning. So what would robot RL look like if we didn’t need to limit compute per interaction? Our latest work, Off-Policy Generative Policy Optimization (OGPO, accepted to ICML26) embarks on answering this question
    Image
    13
  • @servo97
    Sarvesh Patil
    @servo97
    Jul 27
    """once someone launches an attack with an open-weights model whose guardrails they've removed, there will be many more defenders enabled by open-weights models to guard against it."""
    Image
    GIF
    @aran_nayebi
    Aran Nayebi
    @aran_nayebi
    Jul 27
    Anthropic treats attacker access to open models as decisive while ignoring defender access to the same models. That is not a complete offense-defense analysis. Open weights also scale detection, hardening, attribution, and coordinated response. A coalition is the stable
  • @servo97
    Sarvesh Patil
    @servo97
    Jul 8
    Henlo Frens, I'll be presenting OGPO at #ICML2026 tomorrow (9 July) Time ⏰: 2:30 to 4:15 PM KST Location 📍: Hall A, Poster #202 Feel free to stop by for a chat on RL-finetuning, diffusion/flow policies, dexterous manipulation, or existential philosophy! ⛵️
    simchowitzlabpublic.github.io
    OGPO: Sample-Efficient Full-Finetuning of Generative Control Policies
    Off-policy critic for sample efficiency, on-policy PPO over the denoising process for expressivity. The only method that finetunes weak BC policies to near-full success with no expert data in the...
    1
  • @servo97
    Sarvesh Patil
    @servo97
    Jun 17
    Validation loss can be a red herring measure of model inference with naive optimizers. I have observed this since 2016 (🦕). Awesome work on an approach to mitigate this dissonance which further brings better representation learning at test time as an additional benefit!
    @ThomasTCKZhang
    Thomas Zhang
    @ThomasTCKZhang
    Jun 17
    Super excited to finally announce our new paper “Double Preconditioning (DoPr): Optimization for Test-Time Performance, not Validation Loss” Tl;dr: from LLMs to robotics, on-policy deployment causes a mismatch between validation loss over the training distribution and
    Image
  • @servo97
    Sarvesh Patil
    @servo97
    Nov 6, 2023
    Hello all, We are excited to organize the first workshop on Learning for Soft Robots at #CoRL2023! Through our amazing line-up of speakers, we hope to motivate future research in various soft robotic areas such as manipulation, locomotion, sensors, and embodied perception!
    Image
    1
Advertisement
Advertisement