Encoder-decoder is back 😈!!! In DeepSeek V4.1 Flash the first 20 layers build the global KV that the next 20 read from, so prefill costs about half.
Rolled out over the last 24h: throughput doubled, and we are approaching 1T tokens/day on OpenRouter.
30% off to celebrate.
Fast ML inference. Run top AI models using a simple API.



