Pinned
Introducing TRQAM! Internalizing a KL trust region inside the sampling SDE stabilizes off-policy RL fine-tuning of pretrained flow policies. With TRQAM, we lift offline RL success on 50 OGBench tasks from 46% to 68%. 馃У [1/8]
yonghdong.github.io/blog/trqam/

