Thanks @_akhaliq for the tweet.
EMD matches the teacher diffusion model’s marginal distribution by making the old EM great again. Check out our paper at: arxiv.org/abs/2405.16852
EM Distillation for One-step Diffusion Models
While diffusion models can learn complex distributions, sampling requires a computationally expensive iterative process. Existing distillation methods enable efficient sampling, but have notable limitations, such as
📢 Excited to share EM Distillation (EMD), a maximum likelihood method that distills pretrained diffusion models to one-step generators. EMD gracefully interpolates between mode-seeking and mode-covering KL to better capture the teacher's distribution.
arxiv.org/abs/2405.16852