Log inSign up
Simon Matrenok
4 posts
@matrs01

Simon Matrenok

@matrs01
Joined March 2024
56
Following
16
Followers
RepliesRepliesRepostsRepostsMediaMedia

Log in or sign up for X

See what’s happening and join the conversation

Continue with phone
or
Log in with username or email
Terms·Privacy·Cookies·Accessibility·Ads Info·© 2026 X Corp.
  • @matrs01
    Simon Matrenok
    @matrs01
    Jul 14, 2025
    I couldn’t be prouder to share this. 🎉 Our work on Quantile Reward Policy Optimization (QRPO) for LLM RL‑finetuning bridged deep theory and large‑scale practice: * Theory first. We cracked the partition‑function “intractable” myth, reframing it with moment‑generating functions
    @SkanderMoalla
    Skander Moalla
    @SkanderMoalla
    Jul 14, 2025
    🚀 Big time! We can finally do LLM RL fine-tuning with rewards and leverage offline/off-policy data! ❌ You want rewards, but GRPO only works online? ❌ You want offline, but DPO is limited to preferences? ✅ QRPO can do both! 🧵Here's how we do it:
    Image
Advertisement
Advertisement