The king of small models is back! Qwen3.8-27B from @Alibaba_Qwen is open source, and Day-0 support is live in SGLang:
- 206.1 tok/s decode on a single RTX 5090, with our NVFP4 plus DSpark
- 38.28 tok/s decode on DGX Spark
Qwen3.8-27B raises the bar again for what a small model