The Māori language Te Reo has no commercial market to speak of, but a broadcaster in New Zealand's far north has been building speech models for it anyway — and all under a license that keeps the recordings with the communities that originally gave them. Half a world away, in a cellular dead zone in East Africa, farmers point a phone at a cassava leaf and get a diagnosis from a model small enough to live on the handset. Neither project needed permission. And neither could have been rented from a frontier model.
The largest companies on earth are now thinking this way. This month, on the Artificial Analysis Intelligence Index, the strongest closed model scored 61. The strongest open model? 57. The open model came in fourth overall, ahead of models from three of the biggest closed labs. At the end of 2025, about a third of all tokens on OpenRouter were being routed to open-weight models; now the seven highest-volume models on the platform all ship open weights.
In a surprising recent turn, corporate heavyweights are leaning in an encouragingly unexpected direction: On July 24, Mozilla signed an open letter alongside NVIDIA, Microsoft, Meta, IBM, Dell, Mistral, Hugging Face, the Linux Foundation, and Andreessen Horowitz that argued that open weights are too important to be ignored.
We have been here before. Mozilla exists because one company tried to own the front door to the web, and an open community made sure it never could. We bet on open the first time. Open won. Together, we can do it again.
Open weights are where the work happens. A majority of production tokens now route through them, and the seven highest-volume models on OpenRouter are all open weight. Closed models still lead at the frontier, on reasoning and multimodality. However, most production workloads run well below that ceiling. Commodity inputs surrender pricing power. Value moves up to the agentic harness, the layer above the model.
Open model here means weights you can download, run and modify on hardware you control. That category splits. Open weights means the parameters ship under a permissive license with no training code and no data documentation, which describes most of what this report measures. Open source AI, as OSI defines it, additionally requires the training code and enough information about the data to rebuild the system.
K3 debuted first on LMArena's Frontend Code Arena at 1679 Elo, leading in six of seven frontend domains, plus coding, instruction-following and general knowledge.
Frontend Code Arena measures building website and UI code, scored by blind developer votes.
K3 scores 88.3 against Sol's 88.8 on Terminal-Bench 2.1. It wins Program Bench, SpreadsheetBench 2 and BrowseComp, and loses FrontierSWE at 81.2 against Fable 5's 86.6.
Terminal-Bench, Program Bench and BrowseComp measure agent work, running terminal tasks, writing programs and researching the web. FrontierSWE measures resolving real software-engineering tickets.
Fable 5 leads K3 by 92 Elo on GDPval-AA v2, the largest Elo separation among the shared benchmarks, alongside long-context fidelity and conversational polish, which Moonshot concedes still trails.
GDPval-AA v2 measures expert-graded professional knowledge work, where long-context reliability and polish appear.
Data from the Mozilla / SlashData 2026 developer survey. Open models lead in adoption. 79% of developers adding AI functionality use them, against 71% for closed, and the two are largely complementary, with half of developers using both. But production is where teams stall. Only 53% of open-model teams reach production versus 63% for closed. The gap traces to operational tooling and trust.
| Challenge | W. Europe & Israel | N. America | Greater China | South Asia | East Asia ex GC | S. America | E. Europe & CIS | Oceania | All |
|---|---|---|---|---|---|---|---|---|---|
| High infrastructure or compute costs | 25% | 26% | 29% | 28% | 28% | 28% | 29% | 18% | 27% |
| Security, privacy, or compliance concerns | 20% | 27% | 18% | 39% | 29% | 28% | 25% | 22% | 26% |
| Ongoing maintenance and updates | 27% | 26% | 18% | 26% | 20% | 31% | 21% | 25% | 24% |
| Complexity of deployment, hosting, or scaling | 27% | 24% | 19% | 24% | 11% | 30% | 26% | 25% | 23% |
| Lack of specialised support | 17% | 16% | 21% | 31% | 24% | 23% | 23% | 32% | 22% |
| Difficulty evaluating or comparing models | 14% | 17% | 14% | 23% | 16% | 26% | 25% | 18% | 18% |
| Difficulty fine-tuning or customising | 22% | 18% | 18% | 20% | 11% | 22% | 18% | 12% | 18% |
| Difficulty integrating into existing systems | 19% | 21% | 14% | 20% | 7% | 26% | 19% | 20% | 18% |
| Insufficient documentation or learning resources | 18% | 15% | 15% | 17% | 15% | 20% | 24% | 15% | 17% |
| Model performance is not good enough | 18% | 15% | 13% | 22% | 16% | 17% | 19% | 8% | 17% |
| No major challenges | 9% | 21% | 16% | 5% | 14% | 4% | 8% | 12% | 12% |
| Weighted sample size | 286 | 277 | 206 | 192 | 164 | 147 | 98 | 39 | 1411 |
Nine layers and 48 components of the stack scored across 10 criteria (1–5). Click a layer to open its components. Each carries its own criterion scores, maturity grade and open-vs-closed parity verdict, and surfaces some of its most-starred open-source projects.
Hover any cell for detail.
KDA-class linear attention broke runtime compatibility, so loading weights no longer guaranteed serving. Moonshot resolved it upstream by contributing KDA prefix caching to vLLM, shipped with the 27 July weights, which strengthens standardization while concentrating influence over the standard.
Watch: whether the next divergent architecture also lands upstream, or forks the serving layer.
Open-weight AI is a commercial market at multi-hundred-billion-dollar scale, built by funded companies and run in production by global enterprises.
| Company | HQ | Layer | Disclosed funding | Valuation | Revenue signal | Leading investors | Stage |
|---|---|---|---|---|---|---|---|
| Databricks | USA | Enterprise platform | — | — | $5.4B run-rate | — | Pre-IPO |
| DeepSeek | China | Frontier open weights | $7.4B | $50B+ | ~$220M ARR | Liang Wenfeng · Tencent · CATL · China National AI Fund | Private |
| Mistral AI | France | Open weights + platform | $3.05B | ~$14B (talks at €20B) | ~$400M ARR, 20× YoY | ASML · a16z · Lightspeed · Nvidia | Private |
| Moonshot AI | China | Open weights (Kimi) | $3.9B | — | — | Meituan/Long-Z · Alibaba · Tencent · HongShan | Private |
| Zhipu AI | China | Open weights (GLM) | Undisclosed | Public | — | Public (HK IPO 2026), prior Alibaba and Tencent | HK IPO 2026 |
| MiniMax | China | Open weights | Undisclosed | Public | — | Public (HK IPO 2026) | HK IPO 2026 |
| Cohere | Canada | Enterprise / on-prem | $1.7B | — | Command A+ open-sourced May 2026 | Radical Ventures · Nvidia · AMD · Schwarz Group | Private |
| Cerebras | USA | Compute | $2.1B | — | — | Fidelity · Atreides · G42 · Tiger Global | Private |
| Reflection AI | USA | Open weights | $2.13B | — | — | Nvidia · Disruptive · Sequoia · Lightspeed · DST Global | Private |
| Together AI | USA | Inference cloud | $1.334B | — | — | Aramco Ventures · General Catalyst · Prosperity7 · Nvidia | Private |
| Hugging Face | USA | Hub | $400M | — | — | Salesforce · Google · Nvidia · IBM | Private |
| LangChain | USA | Harness tooling | $260M | — | 126k+ stars, 60% dev share | IVP · Sequoia · Benchmark · CapitalG | Private |
More than 70 national AI strategies are live. The strategic question is now which layer of the stack a country can own.
Six weeks later the same lever pointed the other way and found nothing to grip. Access can be revoked and restored. A weight release cannot be withdrawn once the files are distributed. The two decisions differ in their reversibility.
You can switch off a model. You cannot switch off a copy already running on a machine you hold. Sovereignty is this same argument at national scale.
Secretary Bessent, 21 July: the government will examine Chinese open-source models for IP theft, and may sanction.
An outright prohibition gets harder to enforce as adoption grows. Every copy already held is outside the reach of the order.
The browser was the user agent of the open web, code on the user's side negotiating with servers on their behalf. That role is being recreated one layer up. Above the model now sits the agentic harness — the orchestration loop, tools, memory, sandboxes, and permission model. It is where production difficulty concentrates, and where the open-vs-closed, owner-vs-renter contest restarts.
Anthropic, February 2026: roughly 24,000 fraudulent accounts generated more than 16M exchanges, 3.4M of them attributed to Moonshot, in violation of terms of service and regional access restrictions. These exchanges concern earlier Claude models. Fable 5 did not ship until 9 June.
The White House has not publicly connected the February activity to K3's training data. Anthropic notes that distillation is a widely used and legitimate training method. The allegation is covert extraction at industrial scale.
Documented, on the record. Anthropic's February disclosure, Kratsios on 22 July, Bessent on 21 July.
Signal, suggestive. Greenblatt finds Claude-identifying responses difficult to explain as random noise.
Proof, absent. No logs and no forensic package. Moonshot denies. Behavioral forensics become possible now that the weights have been released.
K3 disproportionately identifies itself as Claude, a statistically significant distribution that researchers describe as difficult to explain as random noise. Asked what it is, it names a model from Anthropic and never one of the two current releases.
Limit: it names a model that predates the case. Self-identification is a known artifact of training on web text containing Claude outputs.
The highest cross-vendor similarity in the benchmark. The top four cross-vendor pairs are all K3 against Anthropic models.
Limit: Together AI frames this as capability convergence and a cost comparison, and makes no distillation claim.
The highest among all non-Anthropic models, exceeding the similarity between some pairs of Anthropic's own models.
Limit: measured on Kimi-K2 and Sonnet 4.5. It predates the K3 and Fable case and stands as prior-pattern context only.
Commercial frontier APIs refused the forensic work, because their guardrails could not tell an incident responder from an attacker.
Hugging Face ran the forensics on GLM 5.2, open weight and self-hosted. Attacker data and credentials never left its environment.
| Dimension | Thinking Machines · Inkling US · 15 Jul 2026 | Moonshot · Kimi K3 CN · 27 Jul 2026 |
|---|---|---|
| License | Apache 2.0, unambiguous and known at announcement. | Kimi K3 License, custom. K2 was modified-MIT, and that did not settle K3. |
| Weights-to-API order | Weights first. Nothing to gate. | API 16 July, weights 27 July. |
| Serving footprint | ≥600 GB VRAM (NVFP4). A small cluster. | ~1.4 TB native MXFP4, 64+ accelerators. Open, but not runnable by most who hold it. |
| Upstream contribution | Standard architecture. The existing serving stack already runs it. | KDA broke runtime compatibility. Moonshot fixed it by contributing prefix caching to vLLM, and gained influence over the standard. |
| Evidence at launch | Model card with disclosed limitations. | Self-reported benchmarks on its own harness, with deployed-system comparisons in footnotes. |
| Capacity posture | No hosted dependency to strain. | New subscriptions paused 20 July as demand neared capacity. |
| Smaller sibling | Inkling-Small (276B total / 12B active) previewed, weights pending. | None announced. |
Reversible and low-consequence. Fetching a document, querying a database, listing a calendar. These can largely be permitted by default. A bad read costs little and can be repeated safely.
Side effects that are costly or irreversible. Sending a message, spending against a budget, modifying a record, executing a transaction. This is where confirmation, approval thresholds, cost caps and revocation must concentrate.
The unsolved permission problem is a write problem. The harness ecosystem now spans roughly a dozen frameworks, ten harnesses and three peer protocols, yet no portable model defines which writes an agent may perform unattended, which require human approval, and which are forbidden, across an MCP host, an A2A peer, a direct tool invocation and a framework boundary. The protocols hardened the front door and stopped there. MCP's 2025-11-25 specification moved authorization onto OAuth 2.1, and A2A v1.0 standardized signed Agent Cards, but both stop at authentication. Knowing who an agent is says nothing about what it may do.
The human backstop is failing too. CoSAI's MCP threat model lists consent fatigue, the pattern in which users approve the large majority of prompts, as a top-tier threat. Consent fatigue is itself a write-side failure, because the prompts that matter are the ones authorizing action.
Emerging cross-harness architectures are bypassing the framework deadlock by pulling control up to the meta-harness layer. Architectures like Databricks' open-sourced Omnigent move enforcement above the individual agent, applying stateful, contextual policies that track what a session has done and gate the next write accordingly. One policy requires human approval for a code push once an agent has pulled an unverified package. Another enforces cost caps that pause a session after a set spend.
Closed systems still lead in four places. The first is the integrated harness. No open model appears in the verified top tier of the official Terminal-Bench 2.1 board, and even on a neutral scaffold the best open model trails Opus 4.8 by about four points. Behind that harness sits a data flywheel, since usage routed through a lab's own scaffold feeds back into its next model. The second is long-context fidelity at 1M tokens, where Gemini 3 holds 89% multi-needle retrieval against DeepSeek V4-Pro's 41%. The third is turnkey compliance, with SOC 2, HIPAA, and zero data retention available by default. The fourth is accountability, meaning a counterparty the customer can hold liable.
Compliance and accountability are contracting problems. The integrated harness is a tooling problem. Long-context fidelity is a model problem, and closing it is work only the open labs can do.
Each turns on owning the harness, the memory, and the permission model while those layers are still open.
The 3.3% average gap across coding, reasoning and agentic tasks, and open's OpenRouter token share, especially in agentic coding.
Reverses if: token share stalls while the reasoning gap widens.
The Terminal-Bench spread between lab-owned and independent scaffolds, MCP/A2A governance under the AAIF, and the portable permission spec that still doesn't exist.
Reverses if: the lab-harness lead widens, or a closed platform sets the permission standard first.
Open-lab economics (ARR, raises, the Zhipu/MiniMax IPOs) against metered-pricing breakpoints (~2027–28), with sovereign capacity as counterweight.
Reverses if: sovereign funding lapses or open-lab economics fail to scale.
Under active tracking: misuse capability and how easily safety tuning strips from open weights, hard-friction zones (above all synthetic CSAM and NCII), and whether NTIA's monitoring posture holds.
Reverses if: a major misuse event, or a shift from monitoring to restriction.
There is a test you can run for the rest of this. Look at who is seated in the rooms where AI gets decided, and with what status. The day they seat the people who keep AI open, portable, and widely deployed on equal footing, the shift from renting to owning will have happened. The window is open now. It is closing slowly enough to be easy to ignore, and the lease is shorter than it looks. Build with us.
The Māori language Te Reo has no commercial market to speak of, but a broadcaster in New Zealand's far north has been building speech models for it anyway — and all under a license that keeps the recordings with the communities that originally gave them. Half a world away, in a cellular dead zone in East Africa, farmers point a phone at a cassava leaf and get a diagnosis from a model small enough to live on the handset. Neither project needed permission. And neither could have been rented from a frontier model.
The largest companies on earth are now thinking this way. Stripe cut its inference costs 73 percent by hosting and serving 50 million daily calls on open weights. Microsoft is exploring routing its heaviest Copilot workload to Azure-hosted open-weight models — the world's largest software company is now routing around its own partner's meter.
All this capability arrived while people were still viewing the world through a 2023 lens. This month, on the Artificial Analysis Intelligence Index, the strongest closed model scored 61. The strongest open model? 57. The open model came in fourth overall, ahead of models from three of the biggest closed labs. According to Epoch, the distance is only roughly one release cycle behind. At the end of 2025, about a third of all tokens on OpenRouter were being routed to open-weight models; now the seven highest-volume models on the platform all ship open weights. Inference costs at GPT-4 quality have fallen fiftyfold in three years, and the steepest drops came each time open weights landed.
Governments have moved, too. France has committed €109B to AI infrastructure. Portugal has shipped a fully open model for €5.5M. And Canada has put $890M into a sovereign public supercomputer. All told, 47 countries now restrict foreign processing of critical workloads.
However, challenges still lie ahead. That's because open ships easily, but deploys hard. Only 53 percent of open adopters actually reach production, compared to 63 percent for closed. And instead of the gap closing as the company size goes up, it opens. That data says the gap is a tooling and trust problem. What's more, the digital commons is narrowing toward a single origin — and it's not California. Qwen has 942 million downloads to Llama's 476 million. Chinese models serve 46 percent of routed tokens against 35 percent from the US. In July, 29 countries founded WAICO in Shanghai, with no major Western democracy among them. A commons with one supplier stops being a commons.
We got a preview of what this future could be like on June 12, when an export order took down Anthropic's newest model for every user for 19 days. On that day, businesses around the world learned that the off switch belonged to someone else. When the same lever swung at open weights six weeks later, it found nothing to grip. You can revoke access, but you cannot un-release a file that's already on someone's machine.
In a surprising recent turn, corporate heavyweights are leaning in an encouragingly unexpected direction: On July 24, Mozilla signed an open letter alongside NVIDIA, Microsoft, Meta, IBM, Dell, Mistral, Hugging Face, the Linux Foundation, and Andreessen Horowitz that argued that open weights are too important to be ignored.
We have been here before. Mozilla exists because one company tried to own the front door to the web, and an open community made sure it never could. That was 25 years ago. The same play is running again.
This State of Open Source report is how we are making sense of it all. It's our assessment of where the open layer holds and where it gives. Tell us where you agree — or, more importantly, where you think we're wrong.
The builders are already building. A rented future has deeper pockets; an owned one has more hands — millions more. This story ends the same way every time it is told: The many, building in the open, outbuild the few behind walls. We bet on open the first time. Open won. Together, we can do it again.