Can we efficiently and robustly finetune flow matching models with reinforcement learning using differentiable rewards, in an amortized way?
Hint: use optimal control and match your velocity field with value gradients!
Please come by our poster “Value Gradient Guidance for
Assistant Prof @ CUHK-SZ. PhD at @Mila_Quebec & @UMontreal, MS & BS at @GeorgiaTech and ex-visitor at @MPI_IS. Deep Learning/Foundation Models/3D Generation.


