Prior work has shown a counter-intuitive idea that RLVR boosts reasoning using 100% noisy data to a similar level of using clean data. We show that this is false through a more rigorous process of constructing noisy data.
We demonstrate that RLVR with noisy data leads to worse
🤯 We cracked RLVR with... Random Rewards?!
Training Qwen2.5-Math-7B with our Spurious Rewards improved MATH-500 by:
- Random rewards: +21%
- Incorrect rewards: +25%
- (FYI) Ground-truth rewards: + 28.8%
How could this even work⁉️ Here's why: 🧵
Blogpost: tinyurl.com/spurious-rewar…


