Kimi K3 debuts at #3 on DeepSWE.
It's the first open-weights model that delivers frontier-level performance, achieving results similar to Claude Fable and GPT-5.6 Sol.
At 11k employees, our AI costs are going up. Which model & harness should we use to lower cost but also retain great quality?
We didn't want to blindly trust public benchmarks. So we ran a comprehensive evaluation on our tasks, code base, infra. It's been produced by more than