1. X
  2. Rishabh Agarwal
Log inSign up
Rishabh Agarwal
Periodic Labs
1,664 posts
Rishabh Agarwal profile banner
user avatar

Rishabh Agarwal

Periodic Labs
@agarwl_
Reinforcement Learner
agarwl.github.io
Joined May 2016
879
Following
26.4K
Followers
RepliesRepliesMediaMedia

Log in or sign up for X

See what’s happening and join the conversation

Continue with phone
or
Log in with username or email
Terms·Privacy·Cookies·Accessibility·Ads Info·© 2026 X Corp.
  • Pinned
    user avatar
    Rishabh Agarwal
    Periodic Labs
    @agarwl_
    Apr 28
    I gave a talk at ICLR 2026 about how we are scaling RL on frontier LLMs with 1T+ parameters, on experimental data from our physical lab at Periodic! Here's a rough recording of the talk:
    Image
    00:00
  • user avatar
    Rishabh Agarwal
    Periodic Labs
    @agarwl_
    Aug 24
    Mujoco but real
    user avatar
    Linus ✦ Ekenstam
    @LinusEkenstam
    Aug 24
    I did not expect it to be THIS fun watching robot olympics 😂
    Image
    00:00
  • user avatar
    Rishabh Agarwal
    Periodic Labs
    @agarwl_
    Aug 18
    Every follow-up OPD paper should be required to come with a video analogy
    user avatar
    Kasey Zhang
    Osmosis (YC W25)
    @_WEEXIAO
    Aug 17
    how the student model probably feels during OPD
    Image
    00:00
  • user avatar
    Rishabh Agarwal
    Periodic Labs
    @agarwl_
    Aug 13
    Annoying that sucking supervision bits with a straw works so well for LLMs
  • user avatar
    Rishabh Agarwal
    Periodic Labs
    @agarwl_
    Aug 13
    Good blog, makes you think about the empirical observation that cureent RL methods that work for LLMs are *low bias* - value functions trade off variance with bias, and hasn't shown huge gains yet - small bias from trainer-inference mismatch often is catastrophic for scaled up
    beren.io
    How can LLM RL Work Despite Information-Theoretic Inefficiency
    Epistemic Status: Obviously speculative and maybe obvious. The success of RL in LLMs has been puzzling me for a while. People have developed various information-theoretic style arguments by which...
Advertisement
Advertisement