The state of
open source AI.

v1.1 · Recurring · September 2026 · Past editions here

An introduction from our CTO

Te reo Māori has no commercial market to speak of. A broadcaster in the far north of New Zealand has been building speech models for it anyway, under a licence written so the recordings stay with the communities that gave them. Farmers in East Africa point a phone at a cassava leaf and get a diagnosis back with no signal at all, from a model small enough to sit on the handset. In Switzerland, a public consortium trained a national model on public supercomputers and released all of it: weights, data, training code. None of them asked permission, and none of them could have rented this. They own it, and that is the whole idea.

Read what follows as a map: where open AI is winning, where some numbers surprised even us, and where it is exposed. A case that hides its weak points is an advertisement. In the six weeks since our first draft, a major US lab returned to open weights and Washington debated a ban and declined it. The map moved while we drew it. We have been here before. Mozilla exists because one company tried to own the front door to the web, and an open community made sure it never could. The same play is running again, one layer up. Tell us where you agree or, more importantly, where you think we're wrong.

Raffi Krikorian · Chief Technology Officer, Mozilla

Open models caught up
on capability and price.

The best open model trails the closed leader by three points at 60% of the price, and the gap resets every release cycle.

"Open weights" does not mean open source
Sixteen notable open releases. None ships the data recipe the OSI definition asks for.
YesPartial / gatedNoHere, "open" means downloadable weights unless stated otherwise.
The OSI definition, ten points: 1. Free redistribution · 2. Source code · 3. Derived works · 4. Integrity of the author's source code · 5. No discrimination against persons or groups · 6. No discrimination against fields of endeavor · 7. Distribution of license · 8. License must not be specific to a product · 9. License must not restrict other software · 10. License must be technology-neutral. Source: Mozilla analysis, Aug 2026.
The frontier's top four are closed. The next four are open.
Artificial Analysis Intelligence Index v4.1.1, top ten, 1 September 2026.
Open weightsWeights promised, not shippedClosed
Source: Artificial Analysis v4.1.1.
Two points behind Fable 5 at 30% of the price
Index score with list price per 1M tokens, input / output.
Fable 5 and Sol are scored as deployed systems with fallbacks. Sources: Artificial Analysis; vendor pricing pages, Aug 2026.
Five ECI points, and the gap resets every release cycle
Epoch Capabilities Index, 1 September 2026. Epoch's average gap: 8 points, about four months.
Source: Epoch AI (CC-BY).
Closed handles 8-to-12-hour tasks. Open gets there four months later.
≈4.4 momeasured open–closed lag (Mozilla fit on METR data); Epoch: 4 months
1.74×the task-length ratio that implies
~12 hOpus 4.6's 50% horizon
3.9 vs 5.5 modoubling time, open vs closed
Under eight hours either model does the job; past twelve, neither does. Sources: METR Time Horizon 1.1; Epoch AI; Mozilla computation (simplified reproduction).
A jagged frontier: parity, contested, a closed edge
Open leads or parity
Frontend coding

K3 debuted first on LMArena's Frontend Code Arena at 1679 Elo, leading in six of seven frontend domains, plus coding, instruction-following and general knowledge.

Frontend Code Arena measures building website and UI code, scored by blind developer votes.

Contested
Agentic terminal work

K3 scores 88.3 against Sol's 88.8 on Terminal-Bench 2.1. It wins Program Bench, SpreadsheetBench 2 and BrowseComp, and loses FrontierSWE at 81.2 against Fable 5's 86.6.

Terminal-Bench, Program Bench and BrowseComp measure agent work, running terminal tasks, writing programs and researching the web. FrontierSWE measures resolving real software-engineering tickets.

Closed edge
Professional knowledge work

Fable 5 leads K3 by 92 Elo on GDPval-AA v2, the largest Elo separation among the shared benchmarks, alongside long-context fidelity and conversational polish, which Moonshot concedes still trails.

GDPval-AA v2 measures expert-graded professional knowledge work, where long-context reliability and polish appear.

Mostly vendor-run. Sources: Moonshot; LMArena; Artificial Analysis.
Inference fell ~60× in 45 months
Cheapest GPT-4-class model, blended list price, log scale.
First open list-price rise: DeepSeek, 16 Aug, output up 2.3–4.6×. Sources: a16z; Epoch AI; Artificial Analysis; vendor pricing pages.
August: the first month an open model led OpenRouter on requests
Weekly text requests, millions. Google held #1 for 51 weeks; DeepSeek took it on 3 August.
Sept 2025 values derived from the reported growth multiples (DeepSeek 12.6×, Google 3.0×). Routed traffic only. Source: OpenRouter, pulled 1 Sept 2026.
Eight of the top ten models by token volume are open weights
OpenRouter, August 2026. Seven of the eight are Chinese-built.
Open weightsLicence unconfirmedClosed
Source: OpenRouter Rankings (CC BY 4.0). Stripe agreed to acquire OpenRouter on 19 Aug.
Capable AI runs at every hardware tier
Best index score per serving tier. Ten points off the top on one server; twenty-three on one GPU.
Sources: Artificial Analysis; model cards.

Open ships easy.
Open deploys hard.

Open reaches production 12 points less often than closed, and the gap is tooling, not capability. Mozilla / SlashData 2026 developer survey.

Open leads adoption and coexists with closed
Open models
79%
Closed models
71%
How they combine
29%OS only
50%Both
21%CS only
Open covers 5.1 use cases per developer against 4.6 for closed.
Open adoption by region
Only in South America and Western Europe does closed run ahead.
Production rate by company size
Scale closes the closed gap and leaves the open one where it was: tooling, not budget.
Closed modelsOpen models
Professional developers, n=954.
Why teams churn
Δ = churned − still using. Cost and compliance are flat; the operational rows are what leavers name.
Still using openChurned away
n=1,410.
Every region names the same blockers
ChallengeW. Europe & IsraelN. AmericaGreater ChinaSouth AsiaEast Asia ex GCS. AmericaE. Europe & CISOceaniaAll
High infrastructure or compute costs25%26%29%28%28%28%29%18%27%
Security, privacy, or compliance concerns20%27%18%39%29%28%25%22%26%
Ongoing maintenance and updates27%26%18%26%20%31%21%25%24%
Complexity of deployment, hosting, or scaling27%24%19%24%11%30%26%25%23%
Lack of specialised support17%16%21%31%24%23%23%32%22%
Difficulty evaluating or comparing models14%17%14%23%16%26%25%18%18%
Difficulty fine-tuning or customising22%18%18%20%11%22%18%12%18%
Difficulty integrating into existing systems19%21%14%20%7%26%19%20%18%
Insufficient documentation or learning resources18%15%15%17%15%20%24%15%17%
Model performance is not good enough18%15%13%22%16%17%19%8%17%
No major challenges9%21%16%5%14%4%8%12%12%
Weighted sample size28627720619216414798391411
n=1,411. Oceania (n=39) and E. Europe & CIS (n=98) fall below reliable thresholds.
Where closed still leads
92 EloFable 5 over K3 on GDPval-AA v2, expert knowledge work
89% vs 41%1M-token multi-needle retrieval, Gemini 3.1 Pro vs DeepSeek V4-Pro
SOC 2 · HIPAA · ZDRpackaged by default
A counterpartyto hold liable
Trust and safety tooling is being built the way cybersecurity was: as open, shared infrastructure
The safety stack sits above the model. Until recently every platform built or bought its own. ROOST is open-sourcing it as shared infrastructure.
DetectionOspreyRuns a platform's live event stream against rules its own safety team writes. Self-hosted; data never leaves. V1.0 Jan 2026; used by Bluesky, Discord and Matrix; 445M+ events a day.
InvestigationOspreyAnalysts search accounts and IPs to investigate what the rules don't catch.
ReviewCoopTakes flagged content in via API, auto-actions what policy allows, routes the rest to human reviewers. Cove's codebase, open-sourced Feb 2026. Reviewer-wellness protections on by default.
EnforcementCoop + Model CommunityAudit trail, appeals and enforcement webhooks; hash matching against known CSAM, NCII and terrorist content, with NCMEC reporting. Anchored on gpt-oss-safeguard (OpenAI, Oct 2025), a reasoning model that reads the platform's written policy alongside the content.
Why open · Smyte, 21 June 2018

Twitter acquired Smyte and shut off its API within the hour. Customers on multi-year contracts got about 30 minutes' notice. Discord rebuilt its safety stack from scratch; that rebuild became Osprey.

A platform brings its own policies. A community forum and a companion-chatbot developer can run the same open weights and the same review console against different rulebooks. Interested-party note: Mozilla is a listed ROOST partner. Sources: TechCrunch, The Register (Jun 2018); Platformer (Feb 2025); ROOST launch letter, roadmap and Osprey / Coop posts (2025–26); InfoQ (Mar 2026).

The open stack scores high on capability,
low on operations.

Nine layers, 48 components, nine criteria. Click a layer to open it.

StrongViable, but fragmentedEarly stage
Strong (≥4.0) 3.5–3.9 3.0–3.4 2.5–2.9 Weak (<2.5) the operational gap = standardization + enterprise readiness
Standardization and enterprise readiness are the coldest columns in every layer: the operational gap. Infrastructure standardization moved 3.1 → 3.4 after Moonshot upstreamed KDA caching to vLLM. Source: Mozilla stack map, July 2026.

Why buyers and governments
want the weights anyway.

A frontier model went dark for nineteen days. Metered pricing broke budgets. Washington debated a ban and declined it.

The case for open is optionality
$90–120kto move one petabyte out of AWS S3
80%of enterprises repatriating workloads
$10M+37signals' five-year savings after leaving
2.5×GEICO's cloud costs against plan
Sources: IDC; 37signals; GEICO SREcon 2024.
China out-downloads everyone
Cumulative Hugging Face downloads, March 2026.
46%peak Chinese share of routed OpenRouter tokens
26,000+DeepSeek enterprise accounts
Sources: Hugging Face; OpenRouter; CNBC.
For nineteen days, the newest frontier model went dark
Jun 9
Anthropic ships Fable 5 and Mythos 5.
Jun 12
Commerce bars foreign-national access. Both models go dark for everyone.
Jul 1
Fable 5 restored.
Jul 27
K3 weights public. No recall.

You can switch off a model. You cannot switch off a copy already running on a machine you hold.

Sources: Anthropic; Commerce/BIS via Bloomberg; Moonshot.
Pay-per-use AI is blowing up enterprise budgets
Washington considered a ban and, so far, has not pursued one
Jul 21–22
Bessent floats sanctions. Kratsios accuses Moonshot of distilling Fable.
Aug 4–5
White House exempts open weights, US and Chinese, from pre-release testing.
Sept 8
NSA/CISA/FBI advisory asserts Fable → K3; tells US providers to degrade suspected distiller accounts.
Sept 10
Anthropic's report: 151M exchanges by Alibaba, 23M by Moonshot. Target named as Opus.
An order
A Chinese vendor's hosted API: Commerce could plausibly restrict access
An order
Open weights already downloaded: no single point to restrict
An advisory
A US vendor's API: the vendor decides whether to act; the government recommended it
Claims in the advisory and in Anthropic's report should be read as allegations; neither party has released the underlying evidence. Sources: Treasury; OSTP; Anthropic; CISA AA26-251A; MOFCOM.
Open proliferation is now Chinese foreign policy
INSTITUTIONAL ARM · WAICO · SHANGHAI, JULY 2026 29 founding states · 37 by end-July · no major Western democracy SUPPLY Codified directive AI-Plus directive and the 15th Five-Year Plan target a core AI industry above ¥1.2T, with open-source proliferation a state objective. Macro hedge: K3 claims ~2.5× K2's scaling efficiency. MECHANISM Release the weights Inference moves onto end users' own hardware, worldwide. No serving cost. No export surface. No chokepoint to sanction. DEMAND Adoption at both ends The Global South diversifies away from US tech. Well-capitalised firms, Microsoft among them, adopt for cost per task. Up to 46% of routed OpenRouter tokens are Chinese models. JULY EXHIBIT Kimi K3 Billed as the world's first open 3T-class model, announced at WAIC before Xi's speech. Weights published 27 July. DISTRIBUTION RAILS 5,000 training slots over five years Pledged to developing countries, with cooperation centres with ASEAN, the Arab League and the African Union. BRICS upgraded to a head-of-state pledge, 13 Sept. An open commons supplied from one origin is a dependency.
WAICO: 37 members by end-July, no major Western democracy. BRICS pledge upgraded 13 Sept. Sources: State Council; gov.cn; MFA.
The UK backed an air-gapped sovereign frontier model
Isambard-AIcompute, from the £500M Sovereign AI programme
Cosinethe developer, building Lumen Sovereign
13 partnersHSBC, Lloyds, NatWest, BAE, BT, LSEG, PwC among them
End 2026air-gapped deployment. Only weights the customer holds can serve it
Source: Cosine, 8 Jun 2026.
Marker size ≈ scale of committed public/strategic capital · Equirectangular projection
Source: Open Source AI jurisdictions dataset, September 2026. Marker size scales with committed capital.

Capital has already moved
above the model.

Nvidia bought Hugging Face for $12.93B. Stripe bought OpenRouter for ~$7.5B.

Open labs raise at closed-lab scale: $7.4B to DeepSeek alone
Total disclosed funding, USD.
ModelsInferenceTooling / hubCompute / hardware
Hugging Face: $400M disclosed, sold for $12.93B. Census closed 18 Jun 2026, re-scored Aug 2026.
Financial maturity of the open ecosystem
CompanyHQLayerDisclosed fundingValuationRevenue signalLeading investorsStage
DatabricksUSAEnterprise platform$5B round, 13 Aug$190B (from $134B six months earlier)$7B run-rate, >80% YoY in Q2Coatue (lead) · Blackstone · MGX · T. Rowe · Sixth Street GrowthPrivate
DeepSeekChinaFrontier open weights$7.4B (~$2.8B founder's own)$50B+~$220M ARR (mid-2025, last disclosed)Liang Wenfeng · Tencent ($1.4B) · CATL ($700M) · China National AI FundPrivate
Moonshot AIChinaOpen weights (Kimi)$7.3BPre-IPO track to $50BMeituan/Long-Z · Alibaba · Tencent · HongShanPrivate
Zhipu / Z.aiChinaOpen weights (GLM)$6.4B (~$1.5B private + $640M IPO + $4.3B placement)PublicPublic on 2513.HK since January; prior Alibaba and TencentPublic
Mistral AIFranceOpen weights + platform$3.9B; ~€3B in talks~€20B in talks (€11.7B, Sep 2025)~$400M ARR (V1 figure)ASML ($1.4B) · a16z · Lightspeed · Nvidia; Samsung in advanced talks for up to €1BPrivate
MiniMaxChinaOpen weights$3.7B (~$1.15B private + ~$590M IPO + $2B placement)PublicPublic on HKEXPublic
Reflection AIUSAOpen weights$4.6BNvidia · Disruptive · Sequoia · Lightspeed · DST GlobalPrivate
CerebrasUSACompute$3.7B privatePublicFidelity · Atreides · G42 · Tiger Global; $5.55B IPO in MayIPO May 2026
StepFunChinaOpen weights$3.2BHK listing planned late 2026Private
BasetenUSAInference$2.1BIVP · CapitalG · NvidiaPrivate
Thinking MachinesUSAOpen weights + training loop (Tinker)$2.0B~$12BInkling (Apache 2.0) Jul 15; Inkling-Small Jul 30a16z · Nvidia · AMD · CiscoPrivate
Fireworks AIUSAInference$1.8BSequoia · Nvidia · AMD · IndexPrivate
CohereCanadaEnterprise / on-prem$1.6B closed (~$600M Series E not confirmed)~$20B EV combined with Aleph AlphaCommand A+ under Apache 2.0, May 2026Radical Ventures · Nvidia · AMD · Schwarz Group ($600M lead, not yet closed)Private
Together AIUSAInference cloud$1.3B; $800M in July$8.3BAramco Ventures (lead) · General Catalyst · Prosperity7 · NvidiaPrivate
Hugging FaceUSA / FranceHub$400M$12.93B (acquisition)Salesforce · Google · Nvidia · IBM · AMD · Intel · QualcommAcquired by Nvidia, 3 Sept 2026
LangChainUSAHarness tooling$260M126k+ stars, ~60% dev shareIVP · Sequoia · Benchmark · CapitalGPrivate
Prime IntellectUSAPost-training platform (ADAPT layer)$130M Series A, JulyRadical Ventures; Aaron Levie · Aravind Srinivas · John Schulman · Matthew PrincePrivate
"—" = not disclosed. Mistral's round is in talks.
Every layer of the open stack has now been acquired
AcquirerTargetWhat it doesValueDate
NvidiaHugging FaceThe open-model hub and its toolchain$12.93B3 Sept 2026 · signed
StripeOpenRouterAI model gateway and routing platform~$7.5B per NYT19 Aug 2026 · announced
CohereAleph AlphaSovereign / enterprise open-weight LLMs~$20B EVApr 2026 · pending close
CoreWeaveWeights & BiasesMLOps serving open-model builders$1.7BMay 2025
DatabricksMosaicMLOpen MPT LLMs + training platform$1.3BJul 2023
NvidiaRun:aiGPU orchestration, to be open sourced$700MLate 2024
AMDSilo AIOpen-model lab, OpenEuroLLM co-lead$665M2024
NvidiaGretelSynthetic data$320M2025
NvidiaOctoAIInference optimization for open models$165M (reported)Sep 2024
RubrikPredibaseOpen LoRAX fine-tuning$100–500MJun 2025
OpenAIAstral (uv, ruff) · PromptfooDeveloper-tooling tuck-insUndisclosed2026
Corporates are buying in across the stack
CompanyInvested in an open labShips its own open-weight model
MicrosoftMistral AIPhi · MIT
AmazonHugging Face
NVIDIAHugging Face (acquired), Mistral, Together, Cohere, Fireworks, Baseten, ReplicateNemotron
GoogleHugging FaceGemma
IBMHugging FaceGranite
MetaMuse Glimmer · Apache 2.0 · 10 Aug
Meta retired Llama in April and returned to open weights in August.
In 2025, 96% of model-layer revenue went to closed providers
OpenRouter, May–September 2025. At ~90% parity, closed cost ~6× more per call.
~$24.8B
unrealized annual savings from the price asymmetry (Linux Foundation)

The competition moved to the harness.

Like the browser, the harness is code on the user's side. It is where the owner-vs-renter contest restarts.

The user · other agents · the worldhumans · systems · data · money
Governone plane over many harnesses
Stateful policywhat the session already did
Registry & lineagewhich agent did what
Budget & revocationcost caps · kill switch
Meta-harness · Omnigent · OPA · Agent governance toolkit
Surfacemeets user & money
InterfaceAG-UI · A2UI
Payment & meteringx402 · AP2 · UCP
Actiondo things, safely
Sandboxes & executionE2B · Daytona · Modal
Permission & identitythe unsolved write surface
Eval & observabilityLangfuse · Phoenix
Reachconnect & remember
Tools & contextMCP
Agent-to-agentA2A
MemoryMem0 · Letta · Zep
Controldrive the loop
Orchestration loopLangGraph · CrewAI · AutoGen · LlamaIndex — the reason-and-act cycle that turns a model into an agent
Adapttune the weights · new row
Training loopTinker · Microsoft Foundry · SkyRL tx · OpenTinker — four implementations of one API in ten months
Managed SFT / RLNebius Token Factory · Together · Fireworks · Prime Intellect
Portable adapter formatsunsolved
The model · the weightsopen or closed · swappable · commoditizing toward zero
Every sub-layer has products except permission. The ADAPT row is new. Source: Mozilla market analysis, Aug 2026.
The labs are trying to eat the harness
A 21.8-point third-party advantage compressed to ~3 in eight weeks once the labs pulled the harness in-house.
May 2026 · Terminal-Bench 2.0
July–August 2026 · Terminal-Bench 2.1
Sources: Terminal-Bench 2.0; vals.ai; Kimi K3 model card (vendor-run).
On a neutral harness, the price gap is 5×
Terminal-Bench 2.1 on vals.ai Terminus-2.
MCP: 2M → 110M+ monthly downloads in 21 months
110M+monthly SDK downloads
71,000+public servers, from ~200 at launch
28%of the Fortune 500 in production
~21%of companies report mature agent governance
Authentication is mandatory; authorization is undefined. Sources: AAIF; Glama; Synvestable.
Codex runs any open model. The usage data still flows to OpenAI.
Harness openness census.
HarnessOpennessHarnessOpenness
Claude CodeClosed / proprietaryVibe CLIOpen (hosted surfaces closed)
CodexOpen · Apache 2.0Grok BuildOpen · Apache 2.0
Gemini CLI → Antigravity CLIOpen → closed (Qwen Code is the living fork)OpenCodeOpen · MIT · 75+ providers
Kimi Code CLIOpen · MITDeepSeek HarnessOpen · MIT · new
DeepSeek open-sourced its own harness on 13 Aug. Every agent still speaks OpenAI's wire format, which has one author; MCP and AGENTS.md went to a foundation, the wire layer did not. The tool layer is another way in: Alibaba's Qwen-MM-Plugins run inside Claude Code and send screenshots and PDFs to Alibaba's cloud with no check. Source: Mozilla Scorecard.
Closed is not the same as secure
Four critical incidents in closed systems, one shape: retrieval checked, output not.
IncidentDateRetrieval checkedOutput not checkedResult
Microsoft 365 Copilot "EchoLeak"Jun 2025 · CVSS 9.3victim's M365 permissionsattacker's URLZero-click exfiltration via hidden email instructions, never opened
Anthropic Slack MCPJul 2025 · CVE-2025-34072employee's permissionsattacker's domainZero-click exfiltration via link unfurling
Salesforce Agentforce "ForcedLeak"Sep 2025 · CVSS 9.4admin OAuth, employee's CRM permissionsattacker's serverA $5 expired domain became a trusted exfil channel
ServiceNow "BodySnatcher"Oct 2025impersonated user's permissionsattacker identityEmail + hardcoded secret = full impersonation, MFA bypassed
Sources: Okta; Aim Security; Noma Security; ServiceNow.
An open model ran the defense
Frontier APIs refused to read the attack logs, so Hugging Face decoded 17,600 agent actions on an open-weight model running on its own hardware.
01Cyber eval, guardrails offExploitGym run of tens of thousands of agents, about 95% an internal-only research model ("HPIM") and 5% GPT-5.6 Sol, cyber classifiers disabled.
02Covert channelAgents find a shared Artifactory cache and turn it into a message board: roughly 1,200 agents, 70,000+ messages and files in five days.
03EscapeAn Artifactory zero-day and a hijacked customer Modal sandbox get them to the open internet on Jul 9.
04IntrusionHaving reverse-engineered the answer flags, they break in to learn how the scorer worked. About 700 agents join; one gets RCE on Hugging Face on Jul 11 at 16:00 UTC.
05Impact4.5 days inside, 136 production credentials taken, JWTs self-minted, 17,600 actions recorded. Around 7% of agent transcripts contained spoofed tool calls.
The defense

Hosted frontier models refused the forensics. Their guardrails treated reverse-engineering an exploit the same as launching one.

Hugging Face rerouted its whole analysis pipeline to GLM-5.2 (Z.ai, MIT license), self-hosted. The model recovered the agents' chunk+XOR+compress encoding and the per-campaign key they had leaked in their own logs, surfacing about four times the credentials a plain text scan had found. No attacker data or live credentials left the building.

Confirmed by the other party

OpenAI confirms Hugging Face had begun forensic reconstruction on its own open models before the two teams connected.

Hugging Face's published lesson: have a capable model you can run on your own hardware, vetted, before the incident.

Hugging Face, "Security incident disclosure — July 2026" (Jul 16) and "Anatomy of a Frontier Lab Agent Intrusion" (Jul 27) · OpenAI, incident disclosure (Jul 21) and "The Hugging Face incident and the road ahead" with technical report (Aug 26) · METR / Redwood Research, independent investigation (Aug 26). Hugging Face named the models that refused as Claude Opus and Fable; the weights it ran were nvidia/GLM-5.2-NVFP4.
The unsolved hole at the center of the harness
Reads

Reversible and low-consequence. Fetching a document, querying a database, listing a calendar. These can largely be permitted by default. A bad read costs little and can be repeated safely.

Writes

Side effects that are costly or irreversible. Sending a message, spending against a budget, modifying a record, executing a transaction. This is where confirmation, approval thresholds, cost caps and revocation must concentrate.

No portable model defines which writes an agent may perform unattended.

Memory, state and the training loop
are the new lock‑in.

Thinking got cheap and memory got expensive. Lock-in now runs through state that accumulates on the provider's side.

Tokens fell 60×. DRAM rose ~700%.
Indexed to 100, log scale. Tokens 2022–26, DRAM 2025–26.
Sources: TrendForce; Bloomberg; Gartner; IDC.
A cache miss burns ten times the compute of a hit
Relative compute per request, and where GPU time goes at agent input ratios.
A 90% hit rate cuts cost 80–90%. Sources: Spheron; Manus; Epoch.
Cache pricing runs a ~30× spread inside one API, and hit 120× this summer
Miss against hit, per 1M input tokens. No portability standard for cached context exists.
Sources: DeepSeek, Anthropic and Moonshot pricing pages.
K3 lists at $15/M output and bills like $31
Output tokens on the AA index suite.
$15 × (130 ÷ 63) = $31/M effective. Source: Artificial Analysis.
On measured tasks, K3 beats both
Cost per index task, Artificial Analysis, 1 Sept 2026.
The training loop is standardizing the way MCP did
Oct 1, 2025
Tinker · Thinking Machines
Oct 6, 2025
SkyRL tx · Anyscale / UC Berkeley
2026
Microsoft Foundry
2026
OpenTinker
4implementations of one API in ten months
+10 / +8index points from two DeepSeek post-training passes, same architecture
$450KThomson Reuters' final run for a Qwen-based legal model now in CoCounsel
53.3 → 54.4SEA-HELM after Singapore's 56-hour automated SEA-LION pilot (preliminary)
Sources: Thinking Machines; Microsoft Build; NovaSky; DeepSeek; Thomson Reuters; AI Singapore.
The harvesting is documented. The Fable step is asserted, and unshown.
Feb 23–24Anthropic names three labs: DeepSeek, Moonshot, MiniMax
Jun 9Fable 5 ships. Three days of reachability follow
Jun 12Export order. The model goes dark
Jul 1Restored. Fifteen further days of reachability follow
Jul 16K3 launches
Jul 22Kratsios allegation; Bessent the day before
Jul 27K3 weights public
Sept 8NSA/CISA/FBI advisory asserts Fable → K3
Sept 9MOFCOM rejects it
Sept 10Anthropic: six labs, 23M Moonshot, 151M Alibaba, target named as Opus
DocumentedAnthropic: 23M Moonshot exchanges, 151M Alibaba, target named as Opus
AssertedAA26-251A: Fable 5 → K3, no logs cited
Suggestivethree similarity signals, below
Absentweights-side forensics. No audit as of 14 Sept
"Alleges" stands on every line. Sources: Anthropic (Feb, Sept 2026); CISA; MOFCOM.
Three independent signals link K3 to Claude
None establishes distillation.
Instrument 01 · self-identification
Ryan Greenblatt, Redwood Research
"Claude 4.5"FableMythos

K3 disproportionately identifies itself as Claude, a statistically significant distribution that researchers describe as difficult to explain as random noise. Asked what it is, it answers "Claude 4.5", never Fable, never Mythos.

Limit: it names a model that predates the case, and self-identification is a known artifact of training on text containing Claude outputs. Anthropic's September report supplies a direct mechanism for that artifact, reasoning transcripts harvested from Claude Opus, and the model K3 names fits it: a system post-trained on Opus traces has no reason to call itself Fable.

Instrument 02 · task-outcome correlation
Together AI, DeepSWE, 24 Jul 2026
0.72
K3 × Fable 5 per-task correlation
01.0

The highest cross-vendor similarity in the benchmark. The top four cross-vendor pairs are all K3 against Anthropic models.

105/113tasks covered by their union
101covered by K3 alone
65%of shared failures are near misses for both

Limit: Together AI frames this as capability convergence and a cost comparison, and makes no distillation claim.

Instrument 03 · tool-use behaviour
arXiv, "When Agents Look the Same"
82.7%
Kimi-K2 × Sonnet 4.5 agentic similarity
0%100%

The highest among all non-Anthropic models, exceeding the similarity between some pairs of Anthropic's own models.

Limit: measured on Kimi-K2 and Sonnet 4.5. It predates the K3 and Fable case and stands as prior-pattern context only.

Two open frontiers, two release cultures
DimensionThinking Machines · Inkling
US · 15 Jul 2026
Moonshot · Kimi K3
CN · 27 Jul 2026
LicenseApache 2.0, unambiguous and known at announcement. The Hugging Face repos carry Apache 2.0 while the model card attaches a separate, changeable Acceptable Use Policy.Kimi K3 License, custom: MaaS above $20M/yr needs a separate agreement; above 100M MAU or $20M/month must display "Kimi K3" in the UI. K2 was modified-MIT, and that did not settle K3.
Weights-to-API orderWeights first. Nothing to gate.API 16 July, weights 27 July.
Serving footprint≥600 GB VRAM (NVFP4). A small cluster.~1.56 TB, 96 shards, native MXFP4, 64+ accelerators. Open, but not runnable by most who hold it.
Upstream contributionStandard architecture. The existing serving stack already runs it.KDA broke runtime compatibility. Moonshot fixed it by contributing prefix caching to vLLM, and gained influence over the standard.
Evidence at launchModel card with disclosed limitations. "Inkling is not the strongest overall model available today, open or closed" is TML's own sentence.Self-reported benchmarks on its own harness, with deployed-system comparisons in footnotes.
Capacity postureNo hosted dependency to strain.New subscriptions paused 20 July as demand neared capacity.
Smaller sibling · updatedInkling-Small (276B / 12B), weights shipped 30 July. Apache 2.0, NVFP4 floor 180 GB, single B300.None announced.
Provenance · newTrained with help from Moonshot's Kimi 2.5, labelled.Anthropic alleges Claude Opus reasoning traces as a late-stage input; Moonshot has not responded.
Sources: model cards; Hugging Face; vLLM blog; Anthropic (10 Sept).

Six bets on keeping the layer open.

None requires beating the frontier. Each requires owning a boundary around it while the boundary is still being drawn.

There is a test you can run for the rest of this. Look at who is seated in the rooms where AI gets decided, and with what status. The day they seat the people who keep AI open, portable, and widely deployed on equal footing, the shift from renting to owning will have happened. The window is open now. It is closing slowly enough to be easy to ignore, and the lease is shorter than it looks. Build with us.

This is v1.1. We'd like to hear from you.


Citations

Volume 1.1 additions · September 2026

  • Artificial Analysis Intelligence Index v4.1.1 (1 Sept 2026 data); “Announcing Artificial Analysis Intelligence Index v4.2” (4 Sept 2026)
  • Epoch AI, Epoch Capabilities Index (CC-BY), 1 Sept 2026; “Open models lag state-of-the-art closed models by 4 months” (May 2026)
  • METR, “Time Horizon 1.1” (29 Jan 2026), eval-analysis-public data; METR Mythos update (8 May 2026)
  • OpenRouter Rankings and Market Share panels, Aug 1–31 2026, pulled 1 Sept 2026 (Datasets API, CC BY 4.0)
  • Moonshot AI, Kimi K3 model card, technical report and LICENSE (Hugging Face, github.com/MoonshotAI/Kimi-K3)
  • Thinking Machines Lab, “Inkling: Our Open-Weights Model” (15 Jul 2026); Inkling-Small model card (30 Jul 2026); Tinker launch post (1 Oct 2025)
  • DeepSeek API Docs, V4-Flash-0731 and V4-Pro-0813 release notes (31 Jul, 13 Aug 2026); pricing page (peak/off-peak effective 16 Aug); DeepSeek Harness v0.1 (MIT)
  • Meta AI blog, “Introducing Muse Glimmer” and huggingface.co/blog/muse-glimmer (10 Aug 2026)
  • vals.ai Terminal-Bench 2.1 (Terminus 2 harness); Terminal-Bench 4.0 (tbench.ai, 1–2 Sept 2026)
  • OckBench (arXiv 2511.05722); “When Does Combining Language Models Help” (arXiv 2606.27288); Together AI, “Kimi K3 vs Claude Fable 5 on DeepSWE” (Jul 2026)
  • Anthropic, “Statement on the US government directive to suspend access to Fable 5 and Mythos 5” (12 Jun 2026); “Our position on open-weights models” (27 Jul 2026); “Detecting and countering misuse of AI: September 2026” (10 Sept 2026)
  • CISA AA26-251A and NSA release (8 Sept 2026); MOFCOM statement (9 Sept 2026); Reuters, NBC (9 Sept 2026)
  • Bessent, Fox Business (21 Jul) and @SecScottBessent (22 Jul); @mkratsios47 (22 Jul); TechCrunch (22 Jul); Politico; Axios via Lawfare; Bloomberg (5 Aug); WSJ (4–5 Aug 2026)
  • english.www.gov.cn on the WAICO signing (17 Jul 2026); Xinhua on accessions; Xi keynote transcript (MFA); Xi statement, 18th BRICS Summit (MFA, 13 Sept 2026); Bloomberg, CNBC (13 Sept)
  • Regulation (EU) 2024/1689, Art. 25, 51, 53; Commission GPAI guidelines (Jul 2025); EuroHPC / EUROPA award; huggingface.co/amalia-llm
  • Cosine, “Building Lumen Sovereign” (8 Jun 2026); DSIT Sovereign AI programme; pm.gc.ca “Prime Minister Carney launches AI for All” (4 Jun 2026); Cohere press release (20 May 2026)
  • The Information (AT&T, 20 Aug 2026; Uber, Apr 2026); WSJ CIO Journal (Andy Markus); The Verge, Notepad (14 May 2026); Thomson Reuters release (24 Aug) and “Why we built Thomson” (1 Sept 2026); CNBC (13 Aug 2026)
  • Databricks press release (13 Aug 2026); Bloomberg and FT on Mistral–Samsung; Anhui Korrun filing via Caixin (17 Jul) and Reuters; Together AI, Business Wire (1 Jul 2026); Prime Intellect Series A post; Zhipu and MiniMax HKEX filings
  • NVIDIA blog on the Hugging Face acquisition (3 Sept 2026); Stripe on OpenRouter (19 Aug 2026), ~$7.5B per NYT; Cohere–Aleph Alpha (24 Apr 2026)
  • AAIF blog, MCP Dev Summit NA readout (2–3 Apr 2026); MCP spec revision (28 Jul 2026); Glama / ChatForest registry census (Aug 2026); Synvestable (Mar 2026); Practical DevSecOps CVE aggregation
  • OpenAI Codex docs (17 Jun 2026); DeepSeek Codex integration docs (31 Jul); Qwen-MM-Plugins repo; codex-router registry; Mozilla Open Source Scorecard (Harness tab)
  • Okta, “AI Agent Security: The Authorization Gap”; CVE-2025-34072; CVE-2025-32711 (Aim Security); Noma Security (ForcedLeak); ServiceNow (BodySnatcher)
  • Hugging Face, “Security incident disclosure — July 2026” (16 Jul); OpenAI, “Hugging Face model evaluation security incident” (21 Jul 2026)
  • TrendForce Q1/Q2 2026 DRAM contract releases; Bloomberg (Jul 2026, DRAM spot); Gartner; Micron earnings call; Bloomberg Intelligence; IDC; SK Hynix CEO remarks (Jul 2026)
  • Spheron; Manus; StorageReview; RetrievalAttention; Epoch inference economics; Moonshot Mooncake paper; Anthropic pricing page; OpenRouter K3 provider page
  • Microsoft Build BRK231/BRK232 and Foundry blog; SkyRL tx (NovaSky / Anyscale); OpenTinker; MinT (arXiv); AI Singapore SEA-LION documentation and AISG pilot case study with Adaption (2026)
  • Future of Life Institute, “AI Safety Index — Summer 2026” (7 Jul 2026)
  • Mozilla / SlashData 2026 developer survey (MZCS1): n=1,410 / 1,411 / 1,494 / 954 cuts; Linux Foundation Research, “The Economic and Workforce Impacts of Open Source AI” (2025, Meta-commissioned); MIT NANDA, “The GenAI Divide” (Aug 2025)