Skip to main content

Inference

Per-token Sail API pricing

Core models

USDper 1M tokens
ModelWindowInputCachedOutput
Kimi K3
moonshotai/Kimi-K3
Default (ASAP)2.500.2512.50
Balanced2.000.2010.00
Flex1.250.156.25
GLM-5.3
zai-org/GLM-5.3
Default (ASAP)0.980.183.08
Balanced0.500.122.50
Flex0.400.081.80
GLM-5.3-Flash
zai-org/GLM-5.3-Flash
Default (ASAP)0.110.020.35
Balanced0.080.020.28
Flex0.050.010.18
DeepSeek V4.1 Flash
deepseek-ai/DeepSeek-V4.1-Flash
Default (ASAP)0.150.0060.60
Balanced0.120.0050.48
Flex0.080.0040.30
DeepSeek V4 Pro 0813
deepseek-ai/DeepSeek-V4-Pro-0813
Default (ASAP)0.920.042.77
Balanced0.740.032.22
Flex0.460.021.39
DeepSeek V4 Flash 0731
deepseek-ai/DeepSeek-V4-Flash-0731
Default (ASAP)0.090.020.18
Balanced0.070.020.14
Flex0.050.010.09
Kimi-K2.6
moonshotai/Kimi-K2.6
Default (ASAP)1.000.204.00
Balanced0.450.203.00
Flex0.350.102.00
Gemma 4 31B IT
google/gemma-4-31B-it
Default (ASAP)0.400.200.60
Balanced0.120.080.60
Flex0.060.020.30
Gemma 4 31B IT (NVFP4)
nvidia/Gemma-4-31B-IT-NVFP4
Default (ASAP)0.140.070.40
Balanced0.110.060.32
Flex0.070.040.20
Gemma 4 12B IT
google/gemma-4-12B-it
Default (ASAP)0.300.152.00
Balanced0.100.072.00
Flex0.050.021.00
gpt-oss-120b
openai/gpt-oss-120b
Default (ASAP)0.060.030.40

Flex-only models

Models served with the flex completion window exclusively.
USDper 1M tokens
ModelWindowInputCachedOutput
Qwen3.6 35B A3B
Qwen/Qwen3.6-35B-A3B
Flex0.050.020.40

Notes

  • See Completion Windows for how to use balanced and flex for lower token prices.
    • Not all core models support all windows yet. We regularly bring up new models and expand completion window support for existing ones based on demand. If you have a need that’s not represented above, get in touch.
  • Prompt caching is implicit, based on prefix matching. Optionally, you may use prompt_cache_key as a routing hint to help maximize cache hit rates.
  • See Models for capabilities and other details on supported models.
  • To see what these rates add up to on a full agent workload, use the agent cost calculator.

Sailbox

Notes

  • Usage accrues only while a Sailbox is running.
  • Volume storage is billed separately for each hour the volume exists, even when attached Sailboxes are sleeping, until it’s deleted.
  • For more details, see the Sailbox billing overview.

Pricing Plans

Free

Free

to start, pay-as-you-go

For getting started with Sail

  • $5/month in free credits when you attach a payment method
  • Usage-based pricing with prepaid credits
  • Up to 4 seats
  • Up to 100 concurrent Sailboxes
  • Support viaemailand theSail Research Community Slack

Enterprise

Custom

Enterprise-grade customization and support

  • Volume pricing, billed monthly in arrears
  • HIPAA support with a signed BAA
  • Signed MSA and DPA
  • US-only inference with a 10% surcharge on per-token pricing
  • Full usage history through the API, subject to retention
  • Custom model bringups
  • Uptime and latency SLAs
  • Early access to new features
  • Dedicated support via a private Slack channel
  • Plus all of the additional benefits in Pro

Notes

  • You can upgrade to Pro at any time from your billing page.
  • Get in touch to inquire about an Enterprise plan that suits your needs.