1. X
  2. Archit Sharma
Log inSign up
Archit Sharma
408 posts
Image
user avatar
Archit Sharma
@archit_sharma97
RL, post-training, reasoning research @GoogleDeepMind | co-created: Gemini Deep Think series, DPO | prev: @Stanford @Google Brain @IITKanpur @MILAMontreal
Joined July 2015
372
Following
8,058
Followers
RepliesRepliesMediaMedia

New to X?

Sign up now to get your own personalized timeline!

Create account

By signing up, you agree to the Terms of Service and Privacy Policy, including Cookie Use.

Terms·Privacy·Cookies·Accessibility·Ads Info·© 2026 X Corp.
Don't miss what's happening
People on X are the first to know.
Log inSign up
  • user avatar
    Archit Sharma
    @archit_sharma97
    Dec 1, 2024
    don’t email me unless you have this kind of ambition
    Image
    352K0352K
  • user avatar
    Archit Sharma
    @archit_sharma97
    May 30, 2023
    Ever wondered if the RL in RLHF is really needed? Worried that you might really need to understand how PPO works? Worry no more, Direct Preference Optimization (DPO) allows you to fine-tune LMs directly from preferences via a simple classification loss, no RL required. 🧵 ->
    Image
    GIF
    308K0308K
  • user avatar
    Archit Sharma
    @archit_sharma97
    Dec 12, 2023
    Settled the debate on DPO vs PPO for RLHF today.
    Image
    63K063K
  • user avatar
    Archit Sharma
    @archit_sharma97
    Oct 28, 2024
    I don't have a paper to write this in but there is an interesting property when thinking about iterative RL(HF) algorithms. It seems natural to use an improved policy to sample new data online when training LLMs -- turns out that this just lowers the weight on the KL constraint!
    Image
    63K063K
  • user avatar
    Archit Sharma
    @archit_sharma97
    Jan 11, 2024
    🥺
    user avatar
    Andrew Ng
    @AndrewYNg
    Jan 11, 2024
    It is only rarely that, after reading a research paper, I feel like giving the authors a standing ovation. But I felt that way after finishing Direct Preference Optimization (DPO) by @rm_rafailov @archit_sharma97 @ericmitchellai @StefanoErmon @chrmanning and @chelseabfinn. This
    70K070K
  • user avatar
    Archit Sharma
    @archit_sharma97
    Feb 20, 2024
    High-quality human feedback for RLHF is expensive 💰. AI feedback is emerging as a scalable alternative, but are we using AI feedback effectively? Not yet; RLAIF improves perf *only* when LLMs are SFT'd on a weak teacher. Simple SFT on a strong teacher can outperform RLAIF! 🧵->
    Image
    94K094K
  • user avatar
    Archit Sharma
    @archit_sharma97
    Jul 21, 2025
    I have been waiting for this to be announced, it’s so amazing to see such elegant scaling of the Deep Think system where the same system can now achieve a gold at IMO!
    Image
    Advanced version of Gemini with Deep Think officially achieves gold-medal standard at the Interna...
    From deepmind.google
    22K022K
  • user avatar
    Archit Sharma
    @archit_sharma97
    Aug 1, 2025
    Gemini 2.5 Deep Think is out!! We were able to improve the model substantially since our announcement at I/O, and it is a faster variation of the system that got Gold 🥇at IMO (still getting bronze level performance🥉!!) The model is p good at detailed creative tasks too!
    Image
    Image
    28K028K
  • user avatar
    Archit Sharma
    @archit_sharma97
    Nov 21, 2023
    I cannot emphasize how good Mistral 7B is. It has been crushing some of the experiments I have been doing, makes such a huge difference in the fine-tuned performance. Kudos @MistralAI for creating and open-sourcing this model!
    user avatar
    Percy Liang
    Together AI
    @percyliang
    Nov 21, 2023
    HELM v0.4.0 is out! 1) We have a new frontend (thanks to community contribution from Mike Lay). 2) We have added Mistral 7B, which really is punching above its weight (see crfm.stanford.edu/helm/v0.4.0/#/…), rivaling models an order of magnitude larger on the 16 core scenarios:
    Image
    79K079K
  • user avatar
    Archit Sharma
    @archit_sharma97
    Dec 22, 2021
    Embodied agents such as humans and robots live in a continual non-episodic world. Why do we continue to develop RL algorithms in episodic settings? This discrepancy also presents a practical challenge -- algorithms rely on extrinsic interventions (often humans) to learn ..
    Image
    00:00
  • user avatar
    Archit Sharma
    @archit_sharma97
    May 20, 2025
    definitely some deep deep thoughts
    user avatar
    Sundar Pichai
    Google
    @sundarpichai
    May 20, 2025
    Having a deep think...
    Image
    18K018K
  • user avatar
    Archit Sharma
    @archit_sharma97
    Jan 2, 2024
    This is the “right” way to interpret DPO, ie. setting Z(x) to 1 and removing the DoF. But, I think the narrative of removing RL from RLHF has played its part, so I want to course-correct a bit this new year — is DPO still RL(HF), which requires answering, what is RL? 🧵->
    user avatar
    cider
    @jeffreycider
    Jan 1, 2024
    DPO's method of removing RL from RLHF is so based > previously "forced" to sample trajectories with RL bc direct optimization would require an intractable partition fn Z(x) > observe that the bradley-terry model has a few extra degrees of freedom > simply set Z(x)=1
    Image
    Image
    56K056K
  • user avatar
    Archit Sharma
    @archit_sharma97
    Dec 2, 2023
    So, since we are posting science straight to twitter now, @ericmitchellai and I have some updates for potential overfitting in DPO. TL;DR: we compared DPO to IPO and cDPO (DPO + label noise) on 3 different datasets, and we didn't observe any significant advantage (yet). 🧵->
    Image
    78K078K
  • user avatar
    Archit Sharma
    @archit_sharma97
    May 29, 2020
    I wrote an accessible post on why unsupervised learning is important for reinforcement learning and robotics, and how our work on DADS is improving skill discovery and making it feasible in real life. Check it out here!
    user avatar
    Google AI
    @GoogleAI
    May 29, 2020
    Introducing DADS, a novel unsupervised #ReinforcementLearning algorithm for discovering task-agnostic skills, based on their predictability and diversity, that can be applied to learn a broad range of complex behaviors. Learn more at: goo.gle/2ZMPRxe
    Image
    GIF
Advertisement
Advertisement