Pinned
TailRL has sparked an interesting discussion about moving beyond expected-return RL.
It connects closely to a question we explored in Generative Actor Critic (GAC), posted Dec 2025:
Should policy improvement use a fixed training objective—or can it remain a test-time choice? 🧵



