Pinned
Reinforcement Learning (RL) is the key to aligning diffusion models, but it comes with a curse: Reward Hacking. 🎭
Models often game the proxy reward (e.g., OCR scores) while destroying image quality.
⚡ Introducing GARDO: Reinforcing Diffusion Models without Reward Hacking. 👇


