Log inSign up
Zhoujun (Jorge) Cheng
436 posts
Zhoujun (Jorge) Cheng profile banner
@ChengZhoujun

Zhoujun (Jorge) Cheng

@ChengZhoujun
Ph.D. @UCSanDiego | Scaling RL and agents
San Diego
blankcheng.github.io
Joined November 2021
739
Following
1,289
Followers
RepliesRepliesRepostsRepostsMediaMedia

Log in or sign up for X

See what’s happening and join the conversation

Continue with phone
or
Log in with username or email
Terms·Privacy·Cookies·Accessibility·Ads Info·© 2026 X Corp.
  • Pinned
    @ChengZhoujun
    Zhoujun (Jorge) Cheng
    @ChengZhoujun
    Jan 20
    Pretraining has scaling laws to guide compute allocation. But for RL on LLMs, we lack a practical guide on how to spend compute wisely. We show the optimal compute allocation in LLM RL scales predictably. ↓ Key takeaways below
    Image
    GIF
    18
  • @ChengZhoujun
    Zhoujun (Jorge) Cheng
    @ChengZhoujun
    Jul 8
    Our work on compute-optimal RL is at ICML 2026 this week! — IsoCompute Playbook: Optimally Scaling Sampling Compute for RL Training of LLMs 📍 Poster #1813 @ HALL A, Thu Jul 9, 2:30–4:15pm KST, Seoul. I can't make it this time, but my amazing co-author @TongtongLiang99 will
    @ChengZhoujun
    Zhoujun (Jorge) Cheng
    @ChengZhoujun
    Jan 20
    Pretraining has scaling laws to guide compute allocation. But for RL on LLMs, we lack a practical guide on how to spend compute wisely. We show the optimal compute allocation in LLM RL scales predictably. ↓ Key takeaways below
    Image
    GIF
  • @ChengZhoujun
    Zhoujun (Jorge) Cheng
    @ChengZhoujun
    Jun 26
    I think good benchmarks test what your model can improve; great ones quietly point at where the value is and what's worth improving next. For me the tell: you instantly wonder why it didn't come out sooner, and you want to try it today. OSWorld V2 grounded in such real
    @XLangNLP
    XLANG NLP Lab
    @XLangNLP
    Jun 26
    Two years ago, we built OSWorld 1.0 — the benchmark that became the standard for computer-use agents. Agents now score 83.5% on it. Problem solved? Not even close. 🚀Today we introduce OSWorld 2.0: Benchmarking Computer Use Agents on Long-Horizon Real-World Tasks. What's new:
    Image
    00:00
  • @ChengZhoujun
    Zhoujun (Jorge) Cheng
    @ChengZhoujun
    May 26
    Very solid data work here on cua envs. Massive value in both the outcome and the execution!
    @BowenWangNLP
    Bowen Wang
    @BowenWangNLP
    May 26
    RLVR has become the recipe for agentic post-training. But for Computer-Use Agents, the bottleneck is not the algorithm, it is the data. 🐌 🚀 We introduce CUA-Gym: a scalable, lightweight synthesis engine that turns arbitrary task queries into verifiable RLVR data for
    Image
    00:00
    1
  • @ChengZhoujun
    Zhoujun (Jorge) Cheng
    @ChengZhoujun
    May 19
    Thanks RadixArk for sharing our work NanoRollout!!
    @radixark
    RadixArk
    @radixark
    May 19
    Slow, heavy environments have been the real bottleneck for agentic RL. NanoRollout tackles it head-on with a clean rollout-as-a-service design, integrated with miles for scalable agent RL. Great work from the team!
Advertisement
Advertisement