Releasing the model weights and technical report of Kimi K3.
Kimi K3 is our most capable model: a 2.8T MoE model with native visual understanding and a 1M-token context window.
New model architecture: 2.5x the intelligence per unit of compute, not just more params.
Alongside
Opinions are my own | instagram.com/iimtsuki
- Introducing Kimi K3: Open Frontier Intelligence 🔹 2.8 Trillion Parameters, 1 Million Context, Native Multimodal 🔹 Kimi Delta Attention enables up to 6.3x faster decoding in million-token contexts 🔹 Attention Residuals deliver ~25% higher training efficiency at <2% additional
- 🌘 Kimi-K2.7-Code, our latest coding model, is now released and open-sourced! 🔷 Improved coding & agent performance over K2.6: +21.8% on Kimi Code Bench v2, +11.0% on Program Bench, and +31.5% on MLS Bench Lite. 🔷 Reasoning efficiency: Less overthinking, with 30% lower
- 观后感:1. 所有大厂的 AI PM 都在自嗨,仿佛 iPhone 发布后仍在疯狂雕花功能机的山寨机厂商 2. 无招是傻逼没错,但原作者能写出这七万字的黑话雄文,说明其实你俩也挺般配的😁太牛批了,某大厂内网的万字长文。核心项目组员工,在离职前的长篇吐槽,里面出现好几位高管名字。 从见证 AI 立项,到 300 万日活,最后因管理者、员工、文化矛盾等问题,成为项目最后一位走的人。 drive.google.com/file/d/15q4eat…



