Log inSign up
Shuo Yang
306 posts
@Andy_ShuoYang

Shuo Yang

@Andy_ShuoYang
3rd year phd at Berkeley; Efficient ML System;
Berkeley
andy-yang-1.github.io
Joined February 2023
182
Following
4,944
Followers
RepliesRepliesRepostsRepostsMediaMedia

Log in or sign up for X

See what’s happening and join the conversation

Continue with phone
or
Log in with username or email
Terms·Privacy·Cookies·Accessibility·Ads Info·© 2026 X Corp.
  • Pinned
    @Andy_ShuoYang
    Shuo Yang
    @Andy_ShuoYang
    Aug 21
    Your gaming PC can now serve frontier models at interactive speed using official checkpoints without extreme quantization! Qwen3.6 35B → 8GB RTX 4060 laptop @ 39 tok/s DeepSeek-V4-Flash 284B → RTX 5090 desktop @ 22-25 tok/s GLM-5.2 753B → RTX PRO 6000 workstation @ 15 tok/s
    Image
    00:00
    260
  • @Andy_ShuoYang
    Shuo Yang
    @Andy_ShuoYang
    Sep 3
    Congrats!
    @NunchuxAI
    Nunchux AI
    @NunchuxAI
    Sep 3
    Nunchux is live at nunchux.ai! We’re building the frontier of multimodal generative AI inference: fast, affordable, high-quality serving for image, video, and world models. Our Modelverse brings 30+ image, video, and avatar models behind one API. Use
    Image
  • @Andy_ShuoYang
    Shuo Yang
    @Andy_ShuoYang
    Sep 3
    Qwen3.8-Flash-Next now runs on a single RTX 5090 at 68.3 tok/s — no extreme quantization, no speculative decoding. · Production-level checkpoint: GB300-validated NVFP4 checkpoint from RadixArk · 63GB host RAM — less than what 1-bit quants of this model need · The 51GB n-gram
    Image
    00:00
    37
  • @Andy_ShuoYang
    Shuo Yang
    @Andy_ShuoYang
    Sep 3
    Linear attention is playing a more important role in videogen. Nice work @HaochengXiUCB!
    @HaochengXiUCB
    Haocheng Xi
    @HaochengXiUCB
    Sep 2
    Open-source video generation is now faster than playback without compromising quality. Introducing Video Delta Net (VDN): hybrid attention for live text-to-video with near-lossless quality. VDN accelerates Minimax-H3 by 75 - 90 x, generating 14 seconds of 768p video in 11
    Image
    00:00
    1
  • @Andy_ShuoYang
    Shuo Yang
    @Andy_ShuoYang
    Aug 4
    flashlib + CAKE: geomean 1.9x faster exact KNN on Blackwell 🚀 NVIDIA's CAKE team contributed generated CUDA kernels (tcgen05, per-shape dispatch) — 1.9x geomean / 5.4x peak across 198 shapes vs the previous flashlib release, zero recall loss. Live on flashlib nightly
    Image
    Refresh KNN-search standalone export from the Cake MR 415 head (E2E 1.899x) by yyihuang · Pull...
    From github.com
    4
Advertisement
Advertisement