1. X
  2. SGLang
Log inSign up
SGLang
273 posts
user avatar

SGLang

@sgl_project
Run LLMs fast at any scale πŸ”— github.com/sgl-project/sg… Join our community slack.sglang.io For AI tech blogs & deep-dives πŸ‘‰ @lmsysorg
Palo Alto
sglang.io
Joined May 2025
47
Following
4,559
Followers
RepliesRepliesMediaMedia
  • Pinned
    user avatar
    SGLang
    @sgl_project
    Aug 10
    SGLang v0.5.17 is out πŸš€ Kimi K3 is now in main, optimized with KDA-aware prefix caching, DSpark speculative decoding, and DCP for 1M context. You can run it on both NVIDIA and AMD GPUs, with verified recipes in the SGLang cookbook. We also shipped MiniMax-H3 on
    Image
  • user avatar
    SGLang
    @sgl_project
    7m
    Our ecosystem project and the RL framework behind GLM series training, Slime, just open-sourced its deterministic train–rollout alignment path for GLM-5.2. Megatron training and SGLang rollout matched down to a 4096-token logprob MAE of 1.9e-7, with exact zero hidden-state diff
    user avatar
    slime
    @slime_framework
    16h
    Ever since RL became a core part of LLM training, train–rollout numerical mismatch has been one of the recurring problems in RL infra. There have been several great open-source efforts toward alignment, but production-scale support still often comes with tradeoffs in model
  • user avatar
    SGLang
    @sgl_project
    2h
    Really cool deployment of @NVIDIAAI Nemotron 3.5 Lightning with SGLang + DSpark πŸš€ NVFP4 + speculative decoding + 1M context on DGX Spark / RTX 5090 / RTX 6000 PRO is a great example of how SGLang can make these models practical to run locally. Love seeing the community turn
    user avatar
    Mia
    @MiaAI_lab
    4h
    Run Nemotron 3.5 Lightning on your DGX Spark and RTX 5090/6000 Pro with ease ✨ Snappy model that is great at agentic stuff! Great job @NVIDIAAI πŸ‘ - SGLang - NVFP4 with DSpark - 1M context - 112 tok/s single steam - Up to 362 tok/s at 8 sessions Get it here:
    Image
  • user avatar
    SGLang
    @sgl_project
    3h
    The cache class matrix is dead. Here comes Unified Radix Cache. Prefix caching is the biggest lever in agentic and multi-turn serving, but hybrid models break it: full attention, sliding window, and recurrent state each reuse a different amount of the prefix. Every new
    user avatar
    LMSYS Org
    @lmsysorg
    8h
    πŸš€ New blog: Unified Radix Cache: One Tree for Hybrid Model Prefix Caching Hybrid models complicate prefix caching: each attention type has its own cache reuse semantics, and specialized cache classes multiply combinatorially, duplicating tree logic as caching features grow.
    Image
  • user avatar
    SGLang
    @sgl_project
    9h
    πŸŽ‰ Day 0 support for @nvidia Nemotron 3.5 Lightning is live in SGLang! A 30B hybrid Mamba-Transformer MoE with 3B active params, distilled from Nemotron 3 Ultra for always-on agents: βœ… Three built-in speculators: MTP, DFlash, DSpark βœ… Trained for agent harnesses: coding, tool
    Image
    Image
    user avatar
    NVIDIA AI
    NVIDIA
    @NVIDIAAI
    9h
    Introducing NVIDIA Nemotron 3.5 Lightning⚑ An open 30B MoE model with 3B active parameters, built for always-on agents to complete high-volume, specialized tasks faster. It delivers up to 4x the output speed of similar-sized models.

Log in or sign up for X

See what’s happening and join the conversation

Continue with phone
or
Log in with username or email
TermsΒ·PrivacyΒ·CookiesΒ·AccessibilityΒ·Ads InfoΒ·Β© 2026 X Corp.
Advertisement
Advertisement