🚀New blog: Accelerating Long-Context and Agentic Inference with NVFP4 KV Cache
4-bit KV cache is here! NVFP4 KV in SGLang packs ~1.78× more context into GPU memory and speeds up long-context decode by up to 78%.
Together with @Alibaba_Qwen and @nvidia, we brought NVFP4 KV
🚀 New blog: Cosmos3 in RLinf: 3.33× End-to-End Evaluation Throughput with SGLang
RLinf is an open-source framework designed for embodied intelligence & AI agents. It now supports Cosmos3 from fine-tuning to robot evaluation. With SGLang inference support, we deliver 3.33×
🚀 New Blog: Running DeepSeek-V4-Flash and Kimi-K3 on Consumer Hardware with SSD Expert Pack
These models are far too large for a typical PC's memory. WiCi AI and the SGLang team built SSD Expert Pack: routed experts stay on an NVMe SSD, and the runtime loads only the experts
Thrilled to see day-0 SGLang support for DeepSeek V4.1 Flash!
Technical blog: lmsys.org/blog/2026-09-1…
This is another leap for the DeepSeek family, with several new upgrades to the V4 stack: shared compressed KV, a two-stage sparse indexer, mHC, and Engram.
SGLang supports the
DeepSeek V4.1 Flash weights are out! We are shipping day-0 inference and RL support in SGLang and Miles.
V4.1 extends the V4 stack with compressed KV shared across layers, a two-stage sparse indexer, and a 196B Engram lookup memory.
It is natively multimodal with 552B backbone
SGLang has day-0 support for GLM-5.3! Same runtime, same flags, production-ready on NVIDIA Blackwell and Hopper, and AMD MI300X/325X/355X.
SGLang is also the rollout engine inside Slime, the framework @Zai_org used to post-train GLM-5.3, so the same runtime that generated the RL
GLM-5.3 weights from @Zai_org are live, with SGLang powering day-0 serving support!
GLM-5.3 inherits every optimization and feature we battle-tested for GLM-5.2 over the past months.
On real-world multi-turn agentic workloads, we measured 537.6 tok/s/user on NVFP4 and 413