Great work! Glad to see spectrum preservation serves as a central principle for improving RLVR. We also have proposed a new optimizer called Pion (spherelab.ai/pion/) that is designed with spectrum preservation as the first principle. Pion not only work well for RLVR, but
People keep asking me: what's different about optimization in RL?
Seemingly nothing — the pre-training stack just works (Adam, even SGD 👀 @saagnikkk).
Bringing some answers from my last work (sorry for the delay — been cooking 🚀).
We introduce ISO: Isospectral Optimization:




