Day 0 support for Inkling-Small on Modal.
- 276B parameter MoE with 12B active
- 1M context
- Variable thinking effort
- Native image & audio understanding
- NVFP4 checkpoint fits on a single @NVIDIAAI B300
Great to work with @thinkymachines and @sgl_project on this one.
Today, we are releasing Inkling-Small.
Inkling-Small achieves comparable performance to Inkling at a quarter of its size. It features 276B total parameters, 12B active. We are making the full weights available.
thinkingmachines.ai/news/inkling-s…
Fine-tune it on Tinker today, or chat with
Our Head of Inference, @_gongy , sat down with @cognition's Head of Research, @silasalberti, for a wide-ranging conversation on RL, inference, and the infrastructure work our two teams share, from training frontier coding models to running inference at scale.
Watch here:
I sat down with @silasalberti, Head of Research @cognition, to chat about the intersection of RL and inference -- from training frontier coding models to running inference at scale.
Watch till the end for bonus doggo! :o
0:00 — Intros
1:04 — What's hardest to get right in an RL
Registration is live for Runtime.
Apply to attend our inaugural conference for engineers running AI in production.
October 1st live at The Midway in San Francisco.
Serving a 2.8T model well takes a village.
Grateful to @simon_mo_, @rogerw0108 and everyone at @inferact, who spent five days trading configs and optimizations with us, right up into the early hours of day zero.
And to @Kimi_Moonshot for bringing us together. We're excited to
With Kimi K3 Day-0 on vLLM: Open Frontier Intelligence for Everyone 🚀
At 2.8 trillion parameters, Moonshot AI's Kimi K3 is one of the most powerful open-weight models ever released. Starting today, you can serve it on vLLM the moment the weights are public.
What K3 brings:
🧠