All routes operational

Unify every model one gateway.

One OpenAI-compatible API, CLI and MCP gateway in front of every provider — and every other router. Free models, $2 credit a month, and your own keys at no markup.

Free to start — $2 credit every month, no card. Unlock the shared pool for $1/mo or by donating one working key (earn up to 8% back).

Use it withOpenAIxAIAnthropic+ 175 models across 17 providers
anyrouter ~ openai (python)
import os
from openai import OpenAI

client = OpenAI(
    base_url="https://anyrouter.dev/api/v1",
    api_key=os.environ["ANYROUTER_API_KEY"],
)

resp = client.chat.completions.create(
    model="anthropic/claude-sonnet-4.6",
    messages=[{"role": "user", "content": "Hi"}],
)
ANYROUTER_API_KEYget key

Crowdsourced free tokens

We pool everyone's free & trial keys into one big shared pool — free tokens for everyone. Donate a key, grow the pool, ride free.

01

One endpoint. Every model.

Keep your SDK. One upstream rate-limits — your request doesn't notice.

Point your SDK at one base URL and reference models as provider/model. Every call then runs the same pipeline: authenticate, pick an upstream, retry the next one on failure, cache what repeats, and log every attempt.

Between your call and the model · per request

Key balancing

Spreads load across your keys and quarantines burned ones automatically — no client changes.

Automatic failover

Retries and falls back across providers and pooled keys the instant one errors or rate-limits.

Prompt caching

Reuses cached context across calls to cut repeat token cost and time-to-first-token.

Per-request logs

Every fallback hop, token count and cost captured per attempt — traced and queryable.

03

Everything in one gateway

One key. Every capability you'd otherwise wire up yourself.

No add-ons, no tiers gating the basics. Every capability below ships on the same base URL and the same API key.

04

Why teams switch

Three things that are structurally hard to copy.

Never taxed

Bring your own keys, keep every cent.

Route through your own provider keys and AnyRouter takes no cut — no deposit fee, no per-token markup, no cap. What you pay upstream is exactly what you pay.

Observable

Debug any request.

Every fallback hop, status, latency, token count and cost — captured per attempt. A “debug this request” trace nobody else ships at this tier.

Portable

Your config follows you.

Keys, presets and skills live in AnyRouter, inject locally on demand, and wipe clean on exit. Same setup on your laptop, a server, CI, or a teammate's box.

05

Free models, funded by everyone

Shared key pool

Every member signs up for a provider's free/trial tier — NVIDIA NIM, the Gemini free tier, the Groq free tier, and more — and donates that key. The quotas add up into one pool that every member — you included — can call. Donors earn up to 8% back and their Go plan is free.

06

Unlock Go — two doors, same room

How to join

AnyRouter runs on Go$1/mo, or donate one free-tier provider key (donors ride free and earn up to 8% back). Every Go comes with $2 credit each month.

Create your account

Sign up in seconds — no card required. You land on the dashboard with your first API key ready to copy.

07

The cheapest way in

Three ways to start

A $2 monthly credit and free models come with Go — $1/mo, or free when you donate a provider key. Or bring your own keys at no markup, no card required.

Go plan

$2 credit / month

On Go — $1/mo or a donated key. Here's how far $2 goes on fast models.

Tokens for $2, blended 75% input / 25% output.

Go plan

Free models, 1000/day

Route via anyrouter/free on Go — $0 per token, up to 1000 requests/day.

North Mini Code
$0 in · $0 out
DeepSeek V3.1
$0 in · $0 out
DeepSeek V3.2
$0 in · $0 out
DeepSeek-V4-Flash
$0 in · $0 out
DeepSeek V4 Pro
$0 in · $0 out
Dots3-Note Preview
$0 in · $0 out
Gemma 4 31B
$0 in · $0 out
Ling-3.0-flash
$0 in · $0 out
Ling-3.0-tiny
$0 in · $0 out
Muse Glimmer 30B
$0 in · $0 out
Leanstral 1.5
$0 in · $0 out
Ising Calibration 1.5 31B
$0 in · $0 out
Bring your own keys

Your keys, no markup

11 providers ship a free tier — billed by the provider only, never by us.

OpenCode Zen
Cerebras
Google AI Studio
Mistral
NVIDIA NIM
Ollama Cloud
OpenRouter
QwenCloud
ZenMux
Hugging Face
08

Sign in with AnyRouter

Add AI to your app — your users bring their own AnyRouter

One OAuth button hands your app a scoped, temporary key per user. Every request is billed to that user's own account — no API keys for you to collect, store, or pay for.

Your app's sign-in screen

Click it — see what your users see.

  • OAuth 2.1 + PKCE

    Standard authorization-code flow, public client, no secret to leak. Works from a browser-only app.

  • Inference-only scoped key

    The token can run AI requests and read a basic profile — never keys, billing, or account settings.

  • Per-user billing & revocation

    Each user pays from their own credits or free tier, and can revoke your app any time from their dashboard.

09

175+ models · 17 providers · August 21, 2026

One growing catalog, always current

Ox Alpha joins the catalog free — a stealth reasoning model for coding, agentic work, and long-horizon engineering.

1.0MFree

Ox Alpha is a reasoning model designed for coding, sustained agentic work, and production workloads. It is suited for long-horizon software engineering, complex reasoning, and workflows that combine text with visual context. Ox Alpha is a stealth preview model from a third-party provider who has chosen to remain anonymous during this preview.

128K

Gemma 3 27B IT is Google's instruction-tuned 27B open model for chat, reasoning, and coding, with a 128K context window.

33K

Meta's Llama 3.3 70B instruction-tuned model for high-quality chat, reasoning, and coding at a fraction of frontier cost.

262K

Qwen3.6-27B is Alibaba's flagship dense LLM for agentic coding and complex reasoning. It has a 262K context window (expandable to 1M) and a native thinking mode for repository-scale logic and autonomous programming.

262K

Qwen3.8-27B is Alibaba's dense Qwen3.8 model for agentic coding and complex reasoning, with a 262K context window and native thinking mode.

1.0MZDR

DeepSeek V4 Pro, hosted on Cloudflare Workers AI. Supports a 1,048,576-token context window with function calling and reasoning.

1M

Fast-mode variant of Claude Opus 5 with the same capabilities and a 1,000,000-token context window, tuned for higher output speed.

262K

Seed 2.1 Turbo is a multimodal model from ByteDance Seed for coding and long-horizon agent workflows, with a 262,144-token context window.

262K

Seed 2.0 Code is a ByteDance Seed model optimized for agentic coding, frontend development, and multilingual programming, with a 262,144-token context window.

512K

Dots3-Note Preview is an open-weight mixture-of-experts model from Dots Studio, with 16B active parameters out of 280B total and a 512,000-token context window. Served on the free platform pool.

1.0M

Google's Gemini 3.7 Flash is a multimodal model for fast agentic workflows, coding, and complex multi-step reasoning, with a 1,048,576-token context window.

1.0M

Meta's Muse Spark 1.2 is a multimodal reasoning model for complex agentic tasks. It accepts text, images, video, audio, and PDF documents, returns text, and offers a 1M-token context window.

1.0M

Qwen3.8 2.4T A95B is an open-weight sparse mixture-of-experts model from Qwen, the open-weight variant of Qwen3.8 Max, with 95 billion active parameters out of 2.4 trillion total and a 1,010,000-token context window.

524K

Inkling Small is an open-weight multimodal mixture-of-experts model from Thinking Machines Lab, with 12B active parameters out of 276B total and a 524,288-token context window.

524K

Solar Pro 4 is Upstage's cost-efficient large language model with a 524K context window, built for long-horizon tasks and agentic workflows.

131KFree

Meta Muse Glimmer 30B is a compact multimodal-ready instruct model available on NVIDIA NIM for chat and agent workflows.

131KFree

NVIDIA Nemotron 3.5 Lightning 30B A3B is a sparse MoE chat model optimized for low-latency agentic workloads on NVIDIA NIM.

$2.00 in
$6.00 out
500K

SpaceXAI's smartest model with frontier performance on coding, knowledge work, and STEM. Builds on Grok 4.5 for long-running agents and visual work.

262KFree

Ling 3.0 Tiny is a mixture-of-experts model from inclusionAI, with 1.3B active parameters out of 7.9B total. It is designed for responsive agents, instruction following, and multi-turn conversations, with switchable thinking and instant modes. Served free via the platform free pool, with BYOK as a fallback.

1M

Qwen3.8 Max is Alibaba's flagship 2.4-trillion parameter MoE model from the Qwen3.8 generation with multimodal (text, image, video) support, a 1 million token context window, and top-tier reasoning, coding, and tool use.

262K

Sakana Namazu is Sakana AI's Japanese-specialized LLM, built on the open model Kimi K2.6 and tuned in-house for Japanese and Japanese business contexts. It is a single in-house model (unlike Fugu's orchestrator) offered via an OpenAI-compatible API, and combines web search and code execution to carry complex tasks through to completion.

200K

MiniMax M2 is a large language model for complex software engineering, agentic tool use, and office productivity workflows. Supports complex agent harnesses, dynamic tool search, Agent Teams, and high-fidelity coding and document-editing tasks.

128K

Nemotron Nano 12B V2 VL is a fast, free multimodal model available via NVIDIA's public API. NVIDIA Nemotron Nano 12B V2 is a 12-billion-parameter open multimodal reasoning model designed for video understanding and document intelligence. It introduces a hybrid Transformer-Mamba architecture, combining transformer-level accuracy with Mamba's efficiency. Supports text, image, and video inputs with reasoning capabilities. Available free via the requesty.ai router and BYOK free tier.

128K

Nemotron Nano 9B V2 is a fast, free text model available via NVIDIA's public API. NVIDIA Nemotron Nano 9B V2 is a large language model trained from scratch by NVIDIA, designed as a unified model for both reasoning and non-reasoning tasks. It responds to user queries with controllable reasoning effort and supports structured outputs, tool calling, and 201 languages. Available free via the requesty.ai router and BYOK free tier.

1.0MFree

Ox Alpha is a reasoning model designed for coding, sustained agentic work, and production workloads. It is suited for long-horizon software engineering, complex reasoning, and workflows that combine text with visual context. Ox Alpha is a stealth preview model from a third-party provider who has chosen to remain anonymous during this preview.

128K

Gemma 3 27B IT is Google's instruction-tuned 27B open model for chat, reasoning, and coding, with a 128K context window.

33K

Meta's Llama 3.3 70B instruction-tuned model for high-quality chat, reasoning, and coding at a fraction of frontier cost.

262K

Qwen3.6-27B is Alibaba's flagship dense LLM for agentic coding and complex reasoning. It has a 262K context window (expandable to 1M) and a native thinking mode for repository-scale logic and autonomous programming.

262K

Qwen3.8-27B is Alibaba's dense Qwen3.8 model for agentic coding and complex reasoning, with a 262K context window and native thinking mode.

1.0MZDR

DeepSeek V4 Pro, hosted on Cloudflare Workers AI. Supports a 1,048,576-token context window with function calling and reasoning.

1M

Fast-mode variant of Claude Opus 5 with the same capabilities and a 1,000,000-token context window, tuned for higher output speed.

262K

Seed 2.1 Turbo is a multimodal model from ByteDance Seed for coding and long-horizon agent workflows, with a 262,144-token context window.

262K

Seed 2.0 Code is a ByteDance Seed model optimized for agentic coding, frontend development, and multilingual programming, with a 262,144-token context window.

512K

Dots3-Note Preview is an open-weight mixture-of-experts model from Dots Studio, with 16B active parameters out of 280B total and a 512,000-token context window. Served on the free platform pool.

1.0M

Google's Gemini 3.7 Flash is a multimodal model for fast agentic workflows, coding, and complex multi-step reasoning, with a 1,048,576-token context window.

1.0M

Meta's Muse Spark 1.2 is a multimodal reasoning model for complex agentic tasks. It accepts text, images, video, audio, and PDF documents, returns text, and offers a 1M-token context window.

1.0M

Qwen3.8 2.4T A95B is an open-weight sparse mixture-of-experts model from Qwen, the open-weight variant of Qwen3.8 Max, with 95 billion active parameters out of 2.4 trillion total and a 1,010,000-token context window.

524K

Inkling Small is an open-weight multimodal mixture-of-experts model from Thinking Machines Lab, with 12B active parameters out of 276B total and a 524,288-token context window.

524K

Solar Pro 4 is Upstage's cost-efficient large language model with a 524K context window, built for long-horizon tasks and agentic workflows.

131KFree

Meta Muse Glimmer 30B is a compact multimodal-ready instruct model available on NVIDIA NIM for chat and agent workflows.

131KFree

NVIDIA Nemotron 3.5 Lightning 30B A3B is a sparse MoE chat model optimized for low-latency agentic workloads on NVIDIA NIM.

$2.00 in
$6.00 out
500K

SpaceXAI's smartest model with frontier performance on coding, knowledge work, and STEM. Builds on Grok 4.5 for long-running agents and visual work.

262KFree

Ling 3.0 Tiny is a mixture-of-experts model from inclusionAI, with 1.3B active parameters out of 7.9B total. It is designed for responsive agents, instruction following, and multi-turn conversations, with switchable thinking and instant modes. Served free via the platform free pool, with BYOK as a fallback.

1M

Qwen3.8 Max is Alibaba's flagship 2.4-trillion parameter MoE model from the Qwen3.8 generation with multimodal (text, image, video) support, a 1 million token context window, and top-tier reasoning, coding, and tool use.

262K

Sakana Namazu is Sakana AI's Japanese-specialized LLM, built on the open model Kimi K2.6 and tuned in-house for Japanese and Japanese business contexts. It is a single in-house model (unlike Fugu's orchestrator) offered via an OpenAI-compatible API, and combines web search and code execution to carry complex tasks through to completion.

200K

MiniMax M2 is a large language model for complex software engineering, agentic tool use, and office productivity workflows. Supports complex agent harnesses, dynamic tool search, Agent Teams, and high-fidelity coding and document-editing tasks.

128K

Nemotron Nano 12B V2 VL is a fast, free multimodal model available via NVIDIA's public API. NVIDIA Nemotron Nano 12B V2 is a 12-billion-parameter open multimodal reasoning model designed for video understanding and document intelligence. It introduces a hybrid Transformer-Mamba architecture, combining transformer-level accuracy with Mamba's efficiency. Supports text, image, and video inputs with reasoning capabilities. Available free via the requesty.ai router and BYOK free tier.

128K

Nemotron Nano 9B V2 is a fast, free text model available via NVIDIA's public API. NVIDIA Nemotron Nano 9B V2 is a large language model trained from scratch by NVIDIA, designed as a unified model for both reasoning and non-reasoning tasks. It responds to user queries with controllable reasoning effort and supports structured outputs, tool calling, and 201 languages. Available free via the requesty.ai router and BYOK free tier.