1. X
  2. zhyncs
Log inSign up
zhyncs
Together AI
1,067 posts
Image
user avatar
zhyncs
Together AI
@zhyncs42
LightSeek Mafia 🌁 OPINIONS ARE MY OWN, Senior Director @togethercompute, Governing Board @lightseekorg, Homepage zhyncs.com
Bay Area, CA
zhyncs.com
Joined February 2024
1,073
Following
3,669
Followers
RepliesRepliesMediaMedia

New to X?

Sign up now to get your own personalized timeline!

Create account

By signing up, you agree to the Terms of Service and Privacy Policy, including Cookie Use.

Terms·Privacy·Cookies·Accessibility·Ads Info·© 2026 X Corp.
Don't miss what's happening
People on X are the first to know.
Log inSign up
  • Pinned
    user avatar
    zhyncs
    Together AI
    @zhyncs42
    May 6
    Still amazed by the talent in this community. <2 months after GTC, we rebuilt the stack end-to-end — kernel → scheduler → modeling → frontend — and shipped the fastest open-source MLA attention on Blackwell for agentic workloads. Grateful to be part of it!
    user avatar
    LightSeek Foundation
    @lightseekorg
    May 6
    Introducing TokenSpeed, a speed-of-light LLM inference engine. > TensorRT LLM level performance > vLLM level usability > Built by a lean and mission-driven team in two months > MIT license, open-source github.com/lightseekorg/t… lightseek.org/blog/lightseek…
    Image
    Image
    412K0412K
  • user avatar
    zhyncs
    Together AI
    @zhyncs42
    Mar 2, 2025
    I relocated to San Francisco about a week ago. After getting settled in this week, I'll begin working with the SGLang team @lmsysorg to reproduce the performance achieved by the official DeepSeek team for DeepSeek R1. Please reach out if you're interested in collaborating. Stay
    Image
    35K035K
  • user avatar
    zhyncs
    Together AI
    @zhyncs42
    Aug 11, 2025
    Many companies fork an early version of an inference engine, maintain it internally, and diverge so much that contributing back becomes nearly impossible. slime is different — it always runs on the latest SGLang and continuously contributes improvements back to the community! 🚀
    This post is unavailable.
    27K027K
  • user avatar
    zhyncs
    Together AI
    @zhyncs42
    Mar 11, 2025
    FAKE NEWS
    This post is unavailable.
    32K032K
  • user avatar
    zhyncs
    Together AI
    @zhyncs42
    Jul 6, 2025
    🤯A lot of folks have asked me: how does SGLang iterate so fast? PD disaggregation: ~3 weeks. Large-scale EP: ~1 month. GB200 NVL72 support? Also ~3 weeks. And usually with fewer than 5 core devs involved. SGLang moves fast because of talent like this: ~70k commits in a year!🔥
    Image
    22K022K
  • user avatar
    zhyncs
    Together AI
    @zhyncs42
    Feb 21, 2025
    SGLang will integrate key components from DeepSeek, which will be open-sourced next week to enhance SGLang's inference efficiency. High-performance all-to-all EP implementation is a crucial aspect of this integration!
    user avatar
    Teortaxes▶️ (DeepSeek 推特🐋铁粉 2023 – ∞)
    @teortaxesTex
    Feb 21, 2025
    Interesting guess from SGLang: Expert Parallelism for multi-node inference with DeepSeekMoEs. Maybe we'll finally see their inference economics matched (or really exceeded, if only due to non-gimped hardware)
    55K055K
  • user avatar
    zhyncs
    Together AI
    @zhyncs42
    Aug 23, 2025
    In 2025, open-source LLM serving entered a new era.🚀 SGLang @lmsysorg has been riding the wave with day-0 support and relentless performance gains. In H2, we’re doubling down on stability, usability, and large-scale deployment!👇
    Image
    31K031K
  • user avatar
    zhyncs
    Together AI
    @zhyncs42
    Feb 24, 2025
    "Achieving up to 3000 GB/s in memory-bound configuration and 580 TFLOPS in computation-bound configuration on H800 SXM5, using CUDA 12.6." Great work!
    This post is unavailable.
    20K020K
  • user avatar
    zhyncs
    Together AI
    @zhyncs42
    Aug 2, 2025
    Been in the U.S. for half a year now. Just got my EB1A approved😆 Recently shared the good news with some friends and open-source teammates. It's been a crazy few weeks with all the new model releases. Grateful to amazing collaborators and teammates over the past year🤗
    9.4K09.4K
  • user avatar
    zhyncs
    Together AI
    @zhyncs42
    Dec 26, 2024
    We're excited to announce SGLang @lmsysorg v0.4.1, which now supports DeepSeek @deepseek_ai V3 - currently the strongest open-source LLM, even surpassing GPT-4o.
    Image
    Release Release v0.4.1 · sgl-project/sglang
    From github.com
    19K019K
  • user avatar
    zhyncs
    Together AI
    @zhyncs42
    Feb 24, 2025
    BTW this is a solo work by Jiashi Li. LLM inference is individual heroism.
    user avatar
    zhyncs
    Together AI
    @zhyncs42
    Feb 24, 2025
    "Achieving up to 3000 GB/s in memory-bound configuration and 580 TFLOPS in computation-bound configuration on H800 SXM5, using CUDA 12.6." Great work!
    13K013K
  • user avatar
    zhyncs
    Together AI
    @zhyncs42
    Apr 6, 2025
    Within 12 hours of Llama 4’s release, we optimized its performance to lead other engines by 20%. I’m proud of SGLang’s exceptional engineering and grateful for its vibrant community. More optimizations are on the way—stay tuned!🚀
    user avatar
    LMSYS Org
    @lmsysorg
    Apr 6, 2025
    SGLang now supports Llama 4, the fastest implementation in open source LLM inference engines. It outperforms other engines by 20% when tested on Scout BF16 with H100 TP 8, thanks to the awesome work of Chang Su, @ChengWan17 and @ispobaoke 🚀
    Image
    37K037K
  • user avatar
    zhyncs
    Together AI
    @zhyncs42
    Jun 16, 2025
    SGLang is an early user of FlashInfer and witnessed its rise as the de facto LLM inference kernel library. It won best paper at MLSys 2025, and Zihao now leads its development @NVIDIAAIDev. SGLang’s GB200 NVL72 optimizations were made possible with strong support from the
    user avatar
    LMSYS Org
    @lmsysorg
    Jun 16, 2025
    The SGLang team just ran DeepSeek 671B on NVIDIA’s GB200 NVL72, unlocking 7,583 toks/sec/GPU for decoding w/ PD disaggregation + large-scale expert parallelism — 2.7× faster than H100. Don’t miss this work! 🔥 Thanks to Pen Li from NVIDIA who kicked off this collaboration and
    Image
    11K011K
  • user avatar
    zhyncs
    Together AI
    @zhyncs42
    Feb 25, 2025
    Life update: Thanks to @baseten, @lmsysorg, and many friends, I've moved to San Francisco🌁. This week I am still dealing with some daily chores. If you'd like to meet for coffee in person, schedule a time through my Calendly. Cheers!
    6K06K
Advertisement
Advertisement