misleading paper title, and even the phrasing in the tweet is still a misnomer. this is not about "post-training" but "retrofitting" (cross-architecture distillation)
fwiw i've been bearish on distilling Transformers to recurrent models for a while, they're too different;
Simple beats complicated:
We show that switching to a sliding-window attention mask with attention sinks (at no cost) beats linear attention post-training.
Huge thanks to my collaborators @RheaSukthanker, @CameronPashmina, and @Emy_Aze.
Paper: arxiv.org/abs/2608.28444




