Build with AI for less

The lowest-cost API on the market for image, video, audio, 3D and LLMs. Fully PAYG at the best rates in the industry, no contracts, better rate limits. Access raw GPU compute to run your own inference workloads at 50% lower cost.

Frequently Asked Questions

Coverage and scope

What's on the rate sheet vs the full catalog?
This page is the published pay-as-you-go rate sheet for Runware managed inference. It lists models where Runware has published pricing tiers (currently 300+), across image, video, audio, text, and 3D. Shown rates assume default generation settings unless a row notes otherwise; use the info icon on a row for tier-specific details. This is not every model in the Runware registry. Search also checks the model directory and community registry when priced results are empty. For billing units, tier definitions, and modality-specific rules, see the pricing documentation.
How many models have published pricing on this page?

This page lists models with published pay-as-you-go rates in the Runware content API (currently 300+). The Runware registry supports 400K+ models overall. Use the model directory to explore capabilities, community models, and models without a published rate sheet row.

Why do some directory models not appear here?

Only models with published rates or measured pricing in the Runware pricing API appear in this table. Community models do not have public pricing in the directory today. Check the model detail page or Playground when pricing data exists for a specific model.

Can I see models from the global Diffusion marketplace?

Yes. Marketplace models appear in the model directory. Published rates appear on this page only when Runware has published pricing tiers for that model.

Units and exact cost

How does pricing work?

Runware uses pay-as-you-go billing. Rates on this page are published example prices per generation at the named configuration: images per image, video per generated clip, LLMs per 1M input/output/cache tokens, audio per request, and 3D per generation task. Exact cost depends on your parameters; use Playground to estimate a specific request.

Why do LLM rows show multiple prices?

LLM billing separates input tokens, output tokens, and cache reads. These are different billing dimensions, so collapsed rows show each published rate instead of a single blended number.

How do I estimate what an LLM request costs?

Divide the tokens your request uses by 1,000,000, multiply by the model's per-1M rate, and add the input and output parts together. When a model supports cached input, cache-read tokens are billed at the cached input rate instead of the standard input rate. For example, at GLM-5.3-Flash's base rates of $0.15 input and $0.50 output per 1M tokens, a request with 32k input + 4k output tokens costs 0.032 x $0.15 + 0.004 x $0.50 = $0.0068.

How do I get the exact cost for my request?

Open the model in Playground. Playground estimates cost from your selected parameters. The pricing table shows common published tiers, not every possible request combination.

What kind of image generation costs can I expect?

Image generation typically ranges from fractions of a cent to a few cents per image depending on model, resolution, and quality settings. Expand a model row to see additional tiers.

Savings and billing

What does Save ~N% mean?

When shown, it reflects an active Runware sale on that tier. You will see the original price struck through, the discounted price, and Save ~N% next to it. If there is no strikethrough, there is no savings label.

Does my usage count if I get rate-limited?

No. You are only charged for successful API requests that generate outputs.

Can I get volume discounts?

Yes. High-usage customers can contact sales for custom pricing. See the form below the Serverless section.

Is there a free trial or credit to test with?

Yes. New users who sign up with a business email receive $2 in free credits to explore the platform and test various models.

Serverless and account

How is Serverless compute pricing different?

The Serverless section covers raw GPU and CPU time billed by the second. On-demand model APIs above are managed inference with published per-model tiers. Choose the section that matches your buyer intent.

Are my uploaded models kept private?

Yes. Custom models and content you upload remain private and are not shared or used for training.

Will Runware scale with my needs?

Yes. Infrastructure scales automatically from experiments to production workloads generating millions of assets.