Reference-conditioned diffusion models rely on dense reference token grids. Are they really necessary?
Surprisingly, dropping 80–90% of the reference tokens barely affects quality. A little fine-tuning is enough to recover nearly the same quality at a fraction of the cost.
🌟🚀 Excited to share our latest work: "Keep The Essentials: Efficient Reference Conditioned Generation via Token Dropping"!
TL;DR: Stop wasting compute on redundant tokens! We introduce SparseContext that drops reference tokens for speeding up reference-based image generation⚡






