1. X
  2. Jonathan @RLC
Log inSign up
Jonathan @RLC
285 posts
Jonathan @RLC profile banner
user avatar

Jonathan @RLC

@lightetal
I’m a PhD researcher @MSFTResearch @Caltech @RPI, prev @NECLabsAmerica @AsariAILabs working on LLM-agents, reasoning, RL, test-time scaling, and code generation
San Francisco, CA
jonathanmli.github.io
Joined June 2023
797
Following
800
Followers
RepliesRepliesMediaMedia
  • Pinned
    user avatar
    Jonathan @RLC
    @lightetal
    Feb 26
    Post-training LLMs is like mixing a cocktail: Too much easy data → no learning Too much hard data → instability Wrong balance → collapse And today, we mix it by hand. What if the data mixture could be learned instead of hand-tuned? arxiv.org/abs/2602.20532 🧵👇
    Image
  • user avatar
    Jonathan @RLC
    @lightetal
    Aug 15
    Really excited for my second @RL_Conference ! Looking forward to chatting about RL for increasingly capable agents, self-improvement, curriculum/continual learning, and inference-time search. Would love to meet others thinking about these problems. Say hi at #RLC2026!
  • user avatar
    Jonathan @RLC
    @lightetal
    Aug 12
    Super curious about how this enables collective self-improvement, much like how human society functions today!
    user avatar
    Yisong Yue
    @yisongyue
    Aug 5
    Knowledge flywheels are a new scaling dimension for self-improving AI.
    Article cover image
    Article
    Knowledge Flywheels
    A new scaling dimension is emerging Today, we scale models through better data, architectures, and compute. We scale agents through better tools, search, verification, and harnesses. Tomorrow, we will...
  • user avatar
    Jonathan @RLC
    @lightetal
    Aug 12
    Excited to share that I’m at @MSFTResearch in Montreal until October, working on coding agents and new techniques for agentic training! Hit me up if you’re in the area and want to chat about RL, agents, code generation, post-training, or just grab a drink!
    Image
  • user avatar
    Jonathan @RLC
    @lightetal
    Aug 12
    Amazing work by the team at Asari AI!
    user avatar
    Asari AI
    @AsariAILabs
    Jul 29
    Our self-improving agents optimized the full @vllm_project inference stack, with up to 16% more throughput and interactivity for @deepseek_ai v4 Pro and @Zai_org GLM 5.2 on B200s (no MTP). Every change was verified and our agents got better and faster at it with each iteration.
    Image
    00:00

Log in or sign up for X

See what’s happening and join the conversation

Continue with phone
or
Log in with username or email
Terms·Privacy·Cookies·Accessibility·Ads Info·© 2026 X Corp.
Advertisement
Advertisement