Encoder-decoder is back 😈!!! In DeepSeek V4.1 Flash the first 20 layers build the global KV that the next 20 read from, so prefill costs about half.
Rolled out over the last 24h: throughput doubled, and we are approaching 1T tokens/day on OpenRouter.
30% off to celebrate.
Persimmon is a genuinely different idea: a model of how people actually talk, not another assistant.
Proud to support @humansand on this launch with DeepCluster, a dedicated NVIDIA Blackwell cluster we deploy and operate.
Excited to see where it goes.
deepinfra.com/deepcluster
For AI to work with us, it needs to understand us
Today, we're introducing Persimmon, the first large-scale model designed to realistically simulate how people talk and interact
Two Ling 3.0 flash variants are now live on DeepInfra, both day 0 with @AntLingAGI
→ Ling-3.0-flash-Fin — finance-tuned, 256K context. Source-grounded research across filings and earnings, valuation modeling, spreadsheet-aware output.
→ Ling-3.0-flash-VL — native