Inkling

thinkingmachines/inkling

multimodal MoE reasoning.

cmd --model thinkingmachines/inkling
Intelligence index
40.7
Output speed
56.5 tok/s
Input
$1 /M
Output
$4.05 /M
Cache read
$0.17 /M
Agent-loop cost
$0.42 /M in
Context window
256K tokens
Released
July 15, 2026
Modalities

vs. the lineup

Inkling beside its stablemates and nearest rivals. The ◆ marks the best value in each column across every row shown.

pin a rival:
ModelIntelligenceCodingSpeedInput $/MOutput $/MBlended $/MContext
GPT-5.6 Luna51.271.4184.4$0.10$0.60$0.231.05M
Kimi K2.7 Code41.960.844.4$0.95$4$1.71256K
Tencent Hy341.258.865.2$0.14$0.58$0.25262K
Inkling ◆40.752.156.5$1$4.05$1.76256K
Step 3.7 Flash30.339.6399.5$0.20$1.15$0.44256K
Inkling Small$0.50$1.20$0.681M
DeepSeek V4 Propinned44.359.470.9$0.43$0.87$0.541M
DeepSeek V4 Flashpinned40.356.2122$0.14$0.28$0.181M

coding performance

The Intelligence Index and its sub-scores, ranked against every scored model in the catalog. A metric that has not been measured for Inkling has been left empty.

Coding Index
52.1
#30 of 34 scored
Terminal-Bench
55.1
#30 of 34 scored
Intelligence Index
40.7
#24 of 40 scored
Long-context reasoning
63.3
reasoning across a long context
SciCode
46.1
scientific coding
GPQA Diamond
87.2
graduate-level QA

usage calculator

How far a month of credits goes on Inkling.

Input tokensfresh prompt
800
Output tokensmodel reply
180
Cache read tokensre-read context
50K
cost / request $0.010 · in $1 · out $4.05 · cache $0.17 per M
fresh input 8%output 7%cache reads 85%
Requests / 30 days
997
$10 credits ÷ $0.010 per request
~199 quick fixes~40 bug fixes~7 feature PRs

what real work costs

Real coding tasks priced end to end on Inkling, from a quick lookup to a full-repo agent run.

One agent task
$0.12
180K in at 75% cache hit, 12K out
What you pay
$0.61 /M
all-in across every token that task touched
Sticker input
$1 /M
cache reads bill at $0.17 /M instead
TaskTokens in · outInklingGPT-5.6 LunaClaude Haiku 4.5
Quick lookup / one-liner8K · 1K$0.0061$0.0008$0.0065
Review a 500-line PR60K · 4K$0.04$0.0046$0.04
Fix a bug (agent loop)180K · 12K$0.12$0.01$0.12
Refactor a module320K · 20K$0.21$0.02$0.21
Full-repo agent run900K · 45K$0.50$0.05$0.49

frequently asked

What is Inkling best for?+
Inkling is an open-weight multimodal mixture-of-experts model from Thinking Machines Lab, with 41B active parameters out of 975B total. It is designed for general-purpose reasoning, coding, agentic and tool-use systems, retrieval-augmented generation, instruction following, and multilingual conversational applications. Its native image and audio understanding supports multimodal analysis alongside text.
How much does Inkling cost?+
$1/M input and $4.05/M output, cache reads $0.17/M. In an agent loop most input is cache-read, so the effective input rate is about $0.42/M.
Which plan do I need?+
Available on Go and above.
How do I switch to it?+
Run cmd --model thinkingmachines/inkling, or type /model in a session and pick it. You can switch mid-session without losing context.

Ship code that matches your taste

Command Code is the AI coding agent that continuously learns your taste. Start for $1.

commandcode.ai/models/inkling