I couldn’t be prouder to share this. 🎉
Our work on Quantile Reward Policy Optimization (QRPO) for LLM RL‑finetuning bridged deep theory and large‑scale practice:
* Theory first. We cracked the partition‑function “intractable” myth, reframing it with moment‑generating functions
🚀 Big time! We can finally do LLM RL fine-tuning with rewards and leverage offline/off-policy data!
❌ You want rewards, but GRPO only works online?
❌ You want offline, but DPO is limited to preferences?
✅ QRPO can do both!
🧵Here's how we do it:


