1. X
  2. SGLang
Log inSign up
SGLang
230 posts
user avatar
SGLang
@sgl_project
Run LLMs fast at any scale πŸ”— github.com/sgl-project/sg… Join our community slack.sglang.io For AI tech blogs & deep-dives πŸ‘‰ @lmsysorg
Palo Alto
sglang.io
Joined May 2025
43
Following
4,010
Followers
RepliesRepliesMediaMedia

Log in or sign up for X

See what’s happening and join the conversation

Continue with phone
or
Log in with username or email
TermsΒ·PrivacyΒ·CookiesΒ·AccessibilityΒ·Ads InfoΒ·Β© 2026 X Corp.
  • Pinned
    user avatar
    SGLang
    @sgl_project
    Jul 27
    Day 0 support for Kimi K3 is live in SGLang! The first open 3T-class model: 2.8T params, KDA + Attention Residuals, ultra-sparse MoE, 1M context. Runs across NVIDIA GB300 / B300 / B200 / H200 / H20, AMD MI350X / MI355X, and more. Easter eggπŸ₯š: the whole launch video was made by
    user avatar
    LMSYS Org
    @lmsysorg
    Jul 27
    SGLang day-0 speed on Kimi K3: 423 tok/s (measured on gsm8k), plus RL support ready in Miles @radixark! How the largest open-source model runs this fast: we natively implemented and deeply optimized K3’s new architecture with fused KDA decode kernels, DP attention, DSpark, PD
    Image
    00:00
    3.9K
  • user avatar
    SGLang
    @sgl_project
    5h
    K3 support in SGLang landed for AMD Instinct MI350X and MI355X on the same timeline as everything else. A validated path matters as much as a benchmark number when putting a model into production. Good to see @tensorwave giving developers a place to run it, and we are grateful
    user avatar
    TensorWave
    @tensorwave
    12h
    The open AI ecosystem isn't slowing down, and Kimi K3 (@Kimi_Moonshot) is the latest example. For developers, this release isn't just about benchmark scores. It's about having more choice, greater flexibility, and the freedom to build with the model that best fits the workload.
    936
  • user avatar
    SGLang
    @sgl_project
    12h
    Day 0 SGLang support for Inkling-small and we got 648 tok/s decode with DSpark! Cookbook and HF link below ⬇️
    user avatar
    LMSYS Org
    @lmsysorg
    12h
    Inkling-small is out today! With SGLang, you can get 648 tok/s decode with DSpark (simulated acc len=4) and 288 tok/s w/o DSpark, under the same setup (8x @NVIDIAAI B200, TP 8, NVFP4, bs=1). What makes this model different is the size. 276B total with 12B active is a sweet spot
    Image
    00:00
    1.2K
  • user avatar
    SGLang
    @sgl_project
    12h
    Great to see K3 running fast on Morph, and we're glad to support the team as they scale. Disaggregated serving is where a lot of the remaining headroom in open-model inference sits, and @morphllm is putting it in front of real coding-agent traffic through an OpenAI- and
    user avatar
    Morph
    @morphllm
    14h
    Morph 🀝 SGLang We're working with the @sgl_project and @lmsysorg team to push open-model disaggregated inference faster together. First up is Kimi K3 Fast, served at up to 100 tokens per second through Morph's OpenAI and Anthropic-compatible APIs. This is the beginning of a
    Image
    1.3K
  • user avatar
    SGLang
    @sgl_project
    16h
    Excited to see RadixArk and Google joining the SGLang community to work on TPUs! SGL-JAX is already a production-level solution that delivers fast, native TPU inference for major LLM and diffusion models. And we're looking forward to SGL-torchtpu bringing that same native TPU
    user avatar
    RadixArk
    @radixark
    16h
    RadixArk and Google Cloud are joining forces with the SGLang community to make TPU a drop-in, cost-efficient path to frontier inference. SGL-JAX already serves the major open model families on the latest TPU generations: Gemma, Qwen, DeepSeek, GLM, Kimi, Ling, MiniMax, MiMo,
    Image
    1.9K
  • See @sgl_project's full profile

    Sign up
    Log in
Advertisement
Advertisement