Log inSign up
Yuda Song
338 posts
@yus167

Yuda Song

@yus167
PhD @mldcmu. Previously @ucsd_cse @UcsdMathDept
yudasong.github.io
Joined April 2020
390
Following
1,124
Followers
RepliesRepliesRepostsRepostsMediaMedia

Log in or sign up for X

See what’s happening and join the conversation

Continue with phone
or
Log in with username or email
Terms·Privacy·Cookies·Accessibility·Ads Info·© 2026 X Corp.
  • Pinned
    @yus167
    Yuda Song
    @yus167
    Oct 15, 2025
    🤖 Robots rarely see the true world's state—they operate on partial, noisy visual observations. How should we design algorithms under this partial observability? Should we decide (end-to-end RL) or distill (from a privileged expert)? We study this trade-off in locomotion. 🧵(1/n)
    Image
    2
  • @yus167
    Yuda Song
    @yus167
    Sep 2
    The most common question we get after MaxRL is: how do you handle continuous reward? Now I think we have a good answer to that -- check out the post from @stablegradients to see how we did it 👇 (We'll share more evidence for why we believe this is the correct extension soon.)
    @stablegradients
    Shrinivas Ramasubramanian
    @stablegradients
    Sep 2
    Is RL optimizing the right objective? 🤔 Should we maximize mean reward? Best-of-k? Which k? Standard RL pulls on the mean and often the distribution collapses to a spike. The tail dies 🥲 We introduce Tail-Likelihood Reinforcement Learning (TailRL). It maximizes the mean
    Image
    00:00
    1
  • @yus167
    Yuda Song
    @yus167
    Jul 3
    I will be at ICML 🇰🇷 from Mon to Sat, and present two posters and one oral (more in the thread). Look forward to connecting with friends and discussing RL research (LLMs, robotics, theory), self-improvement, and continual learning.
    Image
    8
  • @yus167
    Yuda Song
    @yus167
    May 19
    Exciting work! But in our February paper, "Reinforcement Learning with Text Feedback", we proposed the same methodology: predicting environment feedback on top of the RL loss. Nice to see this idea specialized to agentic terminal tasks, and the new insight this brings 💡. [1/2]
    Image
    @DimitrisPapail
    Dimitris Papailiopoulos
    @DimitrisPapail
    May 18
    Article cover image
    Article
    ECHO: Terminal Agents Learn World Models for Free
    Co-written with @VaishShrivas We taught CLI agents to predict terminal responses during RL, alongside the usual GRPO loss on actions. The change is tiny: same rollout and forward pass, but stop...
    3
  • @yus167
    Yuda Song
    @yus167
    May 9
    Congrats to my only labmate (in Drew's lab)!
    Image
    @g_k_swamy
    Gokul Swamy
    @g_k_swamy
    Apr 20
    Image
    it took a minute, but i'm proud to share that i'm finally "Dr. Swamy" :)
    1
Advertisement
Advertisement