Astra developed the key ideas for two decade-old conjectures in just two hours. 😱
@OpenAI@sama@gdb
On Sept 6, Xiangxin @NickZhou523786, Jiaxi, and I asked GPT-6-Astra to tackle Chen @wjmzbmr1 & Li’s 2016 best-arm identification conjectures. It produced the key algorithm and
One intuition behind DRPO: the regularizer itself is not the whole story—the trust-region geometry induced by its gradient is what really matters.
By combining advantage-weighted regularization with DPPO-style geometry, DRPO turns hard mask-based trust regions into a smoother
We propose DRPO: a soft version of DPPO🔥
Since PPO, clipping/mask-based trust regions have long outperformed smooth divergence regularization like KL, even though the latter one feels more principled. 👺
We found two missing pieces:👇
1️⃣ Weight the regularizer by |advantage|