Inkling
thinkingmachines/inkling
multimodal MoE reasoning.
cmd --model thinkingmachines/inkling
Intelligence index
42.3
Output speed
75.3 tok/s
Input
$1 /M
Output
$4.05 /M
Cache read
$0.17 /M
Agent-loop cost
$0.42 /M in
Context window
256K tokens
Released
July 15, 2026
Modalities
→
vs. the lineup
Inkling beside its stablemates and nearest rivals. The ◆ marks the best value in each column across every row shown.
pin a rival:
| Model | Intelligence | Coding | Speed | Input $/M | Output $/M | Blended $/M | Context |
|---|---|---|---|---|---|---|---|
| Muse Spark 1.2 Contributor | 56.8◆ | 72.2◆ | —◆ | $0.10◆ | $0.20◆ | $0.13◆ | 1.05M◆ |
| Kimi K2.7 Code | 43◆ | 60.8◆ | 40.5◆ | $0.95◆ | $4◆ | $1.71◆ | 256K◆ |
| MiMo V2.5 Pro | 42.9◆ | 60.2◆ | 51.4◆ | $0.43◆ | $0.87◆ | $0.54◆ | 1M◆ |
| Inkling ◆ | 42.3◆ | 52.1◆ | 75.3◆ | $1◆ | $4.05◆ | $1.76◆ | 256K◆ |
| Inkling Small | 41.2◆ | 52.9◆ | 125.2◆ | $0.50◆ | $1.20◆ | $0.68◆ | 1M◆ |
| Step 3.7 Flash | 30.9◆ | 39.6◆ | 391.3◆ | $0.20◆ | $1.15◆ | $0.44◆ | 256K◆ |
| DeepSeek V4 Pro (latest)pinned | 45.3◆ | 59.4◆ | 63.3◆ | $0.43◆ | $0.87◆ | $0.54◆ | 1M◆ |
| DeepSeek V4 Flash (latest)pinned | 52◆ | 69.1◆ | 115.9◆ | $0.14◆ | $0.28◆ | $0.18◆ | 1M◆ |
coding performance
The Intelligence Index and its sub-scores, ranked against every scored model in the catalog. A metric that has not been measured for Inkling has been left empty.
Coding Index
52.1
#35 of 40 scored
Terminal-Bench
55.1
#34 of 40 scored
Intelligence Index
42.3
#28 of 46 scored
Long-context reasoning
73.3
reasoning across a long context
SciCode
46.1
scientific coding
GPQA Diamond
87.2
graduate-level QA
usage calculator
How far a month of credits goes on Inkling.
Input tokensfresh prompt
800Output tokensmodel reply
180Cache read tokensre-read context
50Kcost / request $0.010 · in $1 · out $4.05 · cache $0.17 per M
fresh input 8%output 7%cache reads 85%
Requests / 30 days
997
$10 credits ÷ $0.010 per request
~199 quick fixes~40 bug fixes~7 feature PRs
what real work costs
Real coding tasks priced end to end on Inkling, from a quick lookup to a full-repo agent run.
One agent task
$0.12
180K in at 75% cache hit, 12K out
What you pay
$0.61 /M
all-in across every token that task touched
Sticker input
$1 /M
cache reads bill at $0.17 /M instead
| Task | Tokens in · out | Inkling | Muse Spark 1.2 Contributor | Claude Haiku 4.5 |
|---|---|---|---|---|
| Quick lookup / one-liner | 8K · 1K | $0.0061 | $0.0003 | $0.0065 |
| Review a 500-line PR | 60K · 4K | $0.04 | $0.0027 | $0.04 |
| Fix a bug (agent loop) | 180K · 12K | $0.12 | $0.0072 | $0.12 |
| Refactor a module | 320K · 20K | $0.21 | $0.01 | $0.21 |
| Full-repo agent run | 900K · 45K | $0.50 | $0.03 | $0.49 |
frequently asked
What is Inkling best for?
Inkling is an open-weight multimodal mixture-of-experts model from Thinking Machines Lab, with 41B active parameters out of 975B total. It is designed for general-purpose reasoning, coding, agentic and tool-use systems, retrieval-augmented generation, instruction following, and multilingual conversational applications. Its native image and audio understanding supports multimodal analysis alongside text.
How much does Inkling cost?
$1/M input and $4.05/M output, cache reads $0.17/M. In an agent loop most input is cache-read, so the effective input rate is about $0.42/M.Which plan do I need?
Available on Go and above.
How do I switch to it?
Run
cmd --model thinkingmachines/inkling, or type /model in a session and pick it. You can switch mid-session without losing context.Ship code that matches your taste
Command Code is the AI coding agent that continuously learns your taste. Start for $1.
Benchmarks from Artificial Analysis (v4.1)commandcode.ai/models/inkling