Pinned
Excited to share Chimera!
As Kimi K3 brings hybrid linear attention into the spotlight, our work explores the same architectural transition from the visual side: adapting KDA-based hybrid attention to token-extensive visual generation, together with a systematic scaling recipe.
Interested in how frontier labs pre-train image/video generation models?
We were too.
Since those recipes are rarely made public in full, we started from the most mature pretraining playbook available in the open: how modern LLMs are built.
Introducing Chimera: a visual






