Review Agent, Coding, and Reasoning model scores, compare provider rates, free tiers, and model coverage, measure API latency, throughput, and duration, and detect model, prompt, and error leakage risks.
Start with model score lookup, then compare per-token pricing, five-round speed tests, health checks, and safety probes before an API reaches production.
Search a model name from the homepage and jump into benchmark pages with AA score, coding, math, and model metadata.
Compare per-token pricing across 100+ providers, find cheaper APIs for each model, and track free tiers and credits.
Run a five-round benchmark with standardized prompts to measure first-token latency, output throughput, and response time.
Audit any OpenAI-compatible API for model authenticity, hidden prompts, instruction tampering, stream integrity, and error leakage, then share a plain-language report.
Enter a base URL, API key, and model ID to test official providers, proxies, relays, or self-hosted endpoints.
Review first-token latency, output throughput, total duration, health, and recent probe signals to judge stability.
How LMSpeed handles model benchmarks, provider pricing, speed tests, and API audits.
Use the model search on the homepage or open the Model Performance leaderboard. Search by model name to find the model detail page, where LMSpeed connects benchmark signals such as AA score, coding, and math with model metadata.
LMSpeed aggregates per-token pricing from 100+ API providers. Visit any model page to see a side-by-side pricing comparison table showing input and output rates per million tokens, so you can find the cheapest provider for each model.
Many providers offer free API tiers or credits for popular models like DeepSeek, Gemini, and Llama. Check our Free LLM API directory for a complete list of models with free access, including speed benchmarks for each free provider.
LMSpeed employs a five-round continuous stress testing mechanism with standardized prompts. Token calculations are performed accurately using tiktoken, measuring output throughput (tokens per second) and first-token latency.
It sends multiple safety probes to an OpenAI-compatible endpoint to check whether the model identity matches, hidden system prompts are injected, user instructions are rewritten, streaming responses stay intact, and errors leak sensitive implementation details. The result includes a risk score and a shareable report. API keys are only used for that audit and are not written to public reports.
Use our performance leaderboards and model detail pages to visually compare API speed benchmarks across providers. The system ranks providers by throughput, latency, and health, helping you choose the fastest and most reliable API.
Provider pages and the health leaderboard already show recent health checks, probe latency, success or failure status, and stability rankings. Broader continuous monitoring and alerting will keep expanding.
A live cut of newly tracked models and benchmark leaders, focused on Artificial Analysis scores for overall intelligence, coding, and math.
| Model | Context | Input | Output | Providers | Agents | Coding | Reasoning | Knowledge | Math | Multilingual | Multimodal | Instruction following | Throughput | Latency | Release date |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Claude Opus 5AnthropicNEW | Context1M | Input$5.00/M | Output$25.00/M | Providers +79 | 65.9±8.2 | 68.4±6.6 | 60.5±8.7 | 67.4±14.0P | — | 64.3±17.1P | 57.6±17.3P | — | Throughput 320 t/s | Latency 4.96s | Release date2026-07-24 |
| Claude Fable 5Anthropic | Context1M | Input$10.00/M | Output$50.00/M | Providers +117 | 64.6±8.2 | 67.3±6.9 | 60.8±10.8E | 59.6±14.0P | — | — | 46.4±17.3P | 50.4±16.0P | Throughput 56 t/s | Latency 3.67s | Release date2026-06-09 |
| Grok 4.6SpaceXAINEW | Context500K | Input$2.00/M | Output$6.00/M | Providers +4 | 62.3±8.6 | 60.5±12.0E | 59.9±10.8E | 59.8±16.0P | — | — | — | — | Throughput 97 t/s | Latency 10.46s | Release date2026-08-12 |
| Kimi K3MoonshotAINEW | Context1.0M | Input$3.00/M | Output$15.00/M | Providers +74 | 68±7.3 | 57.7±12.0E | 59.6±10.8E | 61.6±14.0P | — | — | 66.6±11.3E | — | Throughput 169 t/s | Latency 68.29s | Release date2026-07-16 |
| Qwen3.8 MaxQwenNEW | Context1M | Input$2.00/M | Output$6.00/M | Providers +14 | 66.5±10.9E | 66.1±11.2E | 64.7±10.1E | 52.7±14.0P | — | — | 65.5±6.3 | 57.8±16.0P | Throughput 93 t/s | Latency 3.28s | Release date2026-08-03 |
| Qwen3.8 2.4T A95BQwenNEW | Context1.0M | Input$2.00/M | Output$6.00/M | Providers— | — | 60.4±16.0P | 62.7±13.9P | — | — | — | — | — | Throughput — | Latency — | Release date2026-08-12 |
| GPT-5.6 SolOpenAI | Context1.1M | Input$5.00/M | Output$30.00/M | Providers +131 | 67.2±7.3 | 62.4±9.3E | 61.3±8.7 | 57.2±16.7P | 67±16.1P | — | 59.6±16.1P | 53.5±16.0P | Throughput 57 t/s | Latency 3.65s | Release date2026-07-09 |
| Claude Opus 4.8Anthropic | Context1M | Input$5.00/M | Output$25.00/M | Providers +151 | 62.9±5.5 | 65.6±6.6 | 56.6±8.7 | 64.1±14.0P | 58±16.1P | 55±17.1P | 61±12.0E | 50±16.0P | Throughput 232 t/s | Latency 2.21s | Release date2026-05-27 |
| Muse Spark 1.2MetaNEW | Context1.0M | Input$1.25/M | Output$4.25/M | Providers | 56.9±9.4E | 57.2±12.0E | 61.5±10.8E | 55.8±14.0P | — | — | — | — | Throughput — | Latency — | Release date2026-08-05 |
| GPT-5.5OpenAI | Context1.1M | Input$5.00/M | Output$30.00/M | Providers +146 | 62±5.1 | 59.8±8.7 | 59.3±8.7 | 59.9±14.0P | 56.9±16.1P | — | 55.8±16.1P | 54.7±16.0P | Throughput 46 t/s | Latency 4.95s | Release date2026-04-24 |
| Gemini 3.7 FlashGoogleNEW | Context1.0M | Input$0.750/M | Output$3.75/M | Providers | 63.6±14.1P | 61.7±16.0P | 59.5±13.9P | — | — | — | 59.5±16.1P | — | Throughput — | Latency — | Release date2026-08-13 |
| Grok 4.5SpaceXAI | Context500K | Input$2.00/M | Output$6.00/M | Providers +102 | 61.8±11.3E | 59.8±6.9 | 54.5±8.7 | 58.1±14.0P | — | — | — | — | Throughput 60 t/s | Latency 9.70s | Release date2026-07-08 |
| Claude Sonnet 5Anthropic | Context1M | Input$2.00/M | Output$10.00/M | Providers +107 | 64.8±12.2P | 59.8±6.6 | 56.5±10.8E | 61.4±14.0P | — | — | 58.8±16.1P | — | Throughput — | Latency — | Release date2026-06-30 |
| Claude Opus 4.7Anthropic | Context1M | Input$5.00/M | Output$25.00/M | Providers +195 | 54.6±6.9 | 60.1±11.2E | 56.1±10.8E | 54.6±14.0P | 56.8±16.1P | — | — | 44.5±16.0P | Throughput 47 t/s | Latency 4.89s | Release date2026-05-12 |
| Muse Spark 1.1MetaNEW | Context1.0M | Input$1.25/M | Output$4.25/M | Providers | 63.1±7.4 | 59.3±9.3E | 59.6±10.2E | 65.9±14.0P | — | — | 58.9±16.1P | — | Throughput — | Latency — | Release date2026-07-16 |
| GLM-5.2Z.ai | Context1.0M | Input$1.40/M | Output$4.40/M | Providers +140 | 59.9±7.2 | 54.1±8.9 | 54.7±10.8E | 58.3±14.0P | 70.1±11.5E | — | — | 53.7±16.0P | Throughput 91 t/s | Latency 8.31s | Release date2026-06-16 |
Compare API pricing, speed benchmarks, and performance data across providers.
A unified API gateway providing access to multiple large language models with direct connectivity in China.
Health
100%
Tests
145
Last check
Aug 14
API price
No health checks yet
NVIDIA NIM provides optimized AI model inference APIs for LLMs, vision, and embedding models through NVIDIA cloud infrastructure.
Health
100%
Tests
1,247
Last check
Aug 14
API price
No health checks yet
OpenCode is an open-source AI coding agent that integrates with terminals, IDEs, and desktop apps, supporting multiple models and providers.
Health
100%
Tests
345
Last check
Aug 14
API price
No health checks yet
Suchuang API provides access to various AI models including text generation, image creation, and video generation through a unified API.
Health
100%
Tests
45
Last check
Aug 14
API price
No health checks yet
Ollama provides a platform to run and integrate open-source AI models locally or in the cloud.
Health
100%
Tests
135
Last check
Aug 14
API price
No health checks yet
DeepSeek provides API access to its latest large language models for text generation and coding tasks.
Health
100%
Tests
755
Last check
Aug 14
API price
No health checks yet
The AI gateway built for organizations.
Health
100%
Tests
54
Last check
Aug 14
API price
No health checks yet
VSLLM runs a New API-powered AI gateway on vsllm.com for aggregated model access through a single endpoint.
Health
100%
Tests
155
Last check
Aug 14
API price
No health checks yet
iTokens is an OpenAI-compatible AI API gateway with usage-based access to models from DeepSeek, GLM, Kimi, Qwen, MiniMax, KwaiKAT, and other developers. It publishes per-model token pricing and offers promotional credits to new users.
Volcengine Ark is ByteDance's enterprise AI platform, offering Doubao series models and third-party LLMs via OpenAI-compatible API.
Health
100%
Tests
379
Last check
Aug 14
API price
No health checks yet
Health
100%
Tests
25
Last check
Aug 14
API price
No health checks yet
The newest relay audit reports where endpoint profile, model identity, prompt safety, and response integrity all scored 100.