New - local AI

Kimi K3 now runs locally - Unsloth's 1-bit quant shrunk it 62%

The strongest open model to date (2.8T MoE) fits in RAM on a 748GB DGX Station at 594GB (-62% from 1.56TB), ~78.9% accuracy retained. On a Mac, it loads via SSD memory-mapping with a ~128GB working set - slower, but it runs.

See the model card

Run AI on hardware you own - and follow the models worth running

Track the latest open-weight models, agent harnesses, and autonomous agents, then see which ones fit your rig. Honest speed estimates, cloud-pricing comparisons, and a source-cited tracker of who's running what - so no vendor or government order can switch off the model you depend on.

Latest open-weight models

Browse all models

Latest news & guides

All guides

Top models by quality

Browse all models

Find the right model for your hardware

Already know your rig? Pick it here and see exactly which models you can run locally, with honest speed estimates and cloud-pricing comparisons.

02 - Save your rig (free)

Sign in with GitHub to save your hardware. New here? We'll guide you through picking your rig - Mac, multi-GPU (up to 8x), or custom specs - then show you exactly which models you can run locally.

Sign in with GitHub - free

Who's running what

See all 23 →
Microsoft runs Kimi K3 testing Moonshot AI

The Information: engineers evaluating Moonshot AI's Kimi K3 (2.8T open-weight, released 2026-07-16, $3/$15 per MTok) for Copilot features currently on GPT/Claude, citing strong coding benchmarks and ~60% lower inference cost. Not officially confirmed; evaluating, not deployed.

Microsoft runs MAI reported Microsoft

Bloomberg: Microsoft replacing OpenAI/Anthropic models with in-house MAI models in Excel and Outlook to cut inference spend; a tuned MAI variant claims GPT-5.4 parity at up to 10x efficiency. Proprietary (not open-weight), but the same frontier-API cost pressure.

Smartly runs Llama 3.1 8B confirmed Meta

Self-hosted Llama 3.1 8B on Kubernetes automates support-ticket creation and resolution drafts for the ad-tech platform; 80% less time to create tickets.

Caisse des Depots runs Mistral Medium 3.5 confirmed Mistral AI

Mistral Medium 3.5 (128B) for up to 100k French public-sector agents under a 4-year, EUR 140M framework; on-prem SecNumCloud option for sovereignty.

Capgemini runs Codestral confirmed Mistral AI

Self-hosted Codestral in its RAISE/SovBox coding assistant for regulated aerospace, defense and public-sector clients; code-completion accuracy 50% -> 90%.

Statuses: confirmed official source reported credible third-party, not officially confirmed testing evaluating, not deployed.

Curated and source-cited, not a scraper. Built a rig worth sharing? Browse shared builds →.

41
Models tracked
20
Tools tracked
58
Hardware configs
23
Adopters tracked