Log inSign up
Hexu Zhao
26 posts
Hexu Zhao profile banner
@zhaohexu2001

Hexu Zhao

@zhaohexu2001
Practicing agentic MLSys | building (something) impossible | @NYU_Courant MLSys PhD | @Tsinghua_Uni 23
San Francisco
tarzanzhao.github.io
Joined February 2022
201
Following
233
Followers
RepliesRepliesRepostsRepostsMediaMedia

Log in or sign up for X

See what’s happening and join the conversation

Continue with phone
or
Log in with username or email
Terms·Privacy·Cookies·Accessibility·Ads Info·© 2026 X Corp.
  • @zhaohexu2001
    Hexu Zhao
    @zhaohexu2001
    Aug 31
    Most agentic MLSys work were focused on individual kernels because they’re easy to benchmark, validate, and iterate on. But in production, end-to-end performance is what matters ultimately. Glad to see more work in this direction. I’ve been thinking about this space too, and
    @BrianLi23
    Brian Li
    Baseten
    @BrianLi23
    Aug 28
    Article cover image
    Article
    Agentic Kernels in Production
    TL;DR: We’ve built an agentic kernel development framework that identifies model-level optimization opportunities, generates improved kernels, and validates them in our serving stack. On our current...
  • @zhaohexu2001
    Hexu Zhao
    @zhaohexu2001
    Aug 12
    Claude Code now ships GPU kernels faster than I can review them, using tricks I don't know yet. So my most-used prompt has quietly become: "explain xxx using words I already know." Which is fine once. Less fine when have to ask on every report. 🙃 And the explanation floods my
    1
  • @zhaohexu2001
    Hexu Zhao
    @zhaohexu2001
    Dec 25, 2025
    Our new distributed training research on "point-based differentiable rendering"👇 arxiv.org/abs/2512.20017 We abstract 3 simple APIs🧩to generalize distributed training beyond 3DGS: including static 3D reconstruction (e.g., 3DGS, 2DGS, 3DCX) and dynamic 4D reconstruction (e.g.,
    Image
  • @zhaohexu2001
    Hexu Zhao
    @zhaohexu2001
    Dec 8, 2025
    City-scale 3DGS 🏙️ doesn‘t need a GPU cluster Introducing CLM, a training system for large-scale 3DGS trains 100M+ Gaussian splats on single RTX 4090⚡️ (6x larger than GPU-only while keeping up to 90% throughput) 🔗 Homepage + Code are released! 📄 Accepted to ASPLOS 2026 🧵
    Image
    00:00
    15
Advertisement
Advertisement