1. X
  2. LMSYS Org
Log inSign up
LMSYS Org
1,213 posts
user avatar
LMSYS Org
@lmsysorg
Large Model Systems Organization: We developed SGLang @sgl_project (sglang.io), Chatbot Arena (now @arena), and Vicuna!
US
lmsys.org
Joined August 2024
202
Following
16.5K
Followers
RepliesRepliesMediaMedia

Log in or sign up for X

See what’s happening and join the conversation

Continue with phone
or
Log in with username or email
Terms·Privacy·Cookies·Accessibility·Ads Info·© 2026 X Corp.
  • Pinned
    user avatar
    LMSYS Org
    @lmsysorg
    Jul 27
    SGLang day-0 speed on Kimi K3: 423 tok/s (measured on gsm8k), plus RL support ready in Miles @radixark! How the largest open-source model runs this fast: we natively implemented and deeply optimized K3’s new architecture with fused KDA decode kernels, DP attention, DSpark, PD
    Image
    00:00
    Image
    user avatar
    Kimi.ai
    @Kimi_Moonshot
    Jul 27
    Releasing the model weights and technical report of Kimi K3. Kimi K3 is our most capable model: a 2.8T MoE model with native visual understanding and a 1M-token context window. New model architecture: 2.5x the intelligence per unit of compute, not just more params. Alongside
    192K
  • user avatar
    LMSYS Org
    @lmsysorg
    12h
    Inkling-small is out today! With SGLang, you can get 648 tok/s decode with DSpark (simulated acc len=4) and 288 tok/s w/o DSpark, under the same setup (8x @NVIDIAAI B200, TP 8, NVFP4, bs=1). What makes this model different is the size. 276B total with 12B active is a sweet spot
    Image
    00:00
    user avatar
    Thinking Machines
    @thinkymachines
    12h
    Today, we are releasing Inkling-Small. Inkling-Small achieves comparable performance to Inkling at a quarter of its size. It features 276B total parameters, 12B active. We are making the full weights available. thinkingmachines.ai/news/inkling-s… Fine-tune it on Tinker today, or chat with
    62K
  • user avatar
    LMSYS Org
    @lmsysorg
    Jul 29
    🚀 New blog: Towards Blackwell-Native 8-bit and 4-bit RL: End-to-End MXFP8 and NVFP4 RL in Miles We bring two Blackwell-native, low-precision RL recipes to Miles, fully open-sourced across the stack: ☑️ End-to-end MXFP8 across rollout, forward, and backward GEMMs ☑️
    Image
    14K
  • user avatar
    LMSYS Org
    @lmsysorg
    Jul 27
    LMSYS exists because of open weights. In early 2023, LMSYS started with Vicuna, an open-weight model fine-tuned from Meta's LLaMA (the first phenomenal open-weight model). FastChat, and later SGLang, came out of that same pattern: open weights show up, a community forms around
    Image
    user avatar
    SGLang
    @sgl_project
    Jul 27
    Count us in! SGLang community has co-signed the Open Weights letter @lmsysorg Everything we ship goes out in the open, because world-class performance should be accessible to every builder. Let's keep building in the open 🧡
    21K
  • user avatar
    LMSYS Org
    @lmsysorg
    Jul 25
    SGLang v0.5.16 is out with DSpark, new model support (Inkling, π0.5, LongCat 2.0, and more), and new features!
    user avatar
    SGLang
    @sgl_project
    Jul 25
    🎉 SGLang v0.5.16 is out! This cycle we landed DSpark, a new speculative decoding algorithm. It stays ahead of MTP across the whole concurrency sweep, where speculation usually stops paying off. And we see 383.7 tok/s at accept length ~5 on DeepSeek-V4-Pro (TP8, B300, bs=1). We
    Image
    3K
  • See @lmsysorg's full profile

    Sign up
    Log in
Advertisement
Advertisement