Thanks for the shout out! This work takes quantization to the next level by avoiding straight through gradient estimator for QAT and preventing quality drop. Glad to see it being adopted!
arxiv.org/pdf/2608.13966
Tested the QUASAR-QAT Qwen3.8-27B NVFP4 build on a single DGX Spark last night. vLLM, util 0.37, 131k context.
Plain decode: 12.3 tok/s
With MTP speculation (2 draft tokens): 21.8 tok/s, 82% of drafts accepted
vLLM 0.28 vs 0.26: no difference on either setting
Structured output







