Log inSign up
ali
Baseten
685 posts
ali profile banner
@waterloo_intern

ali

Baseten
@waterloo_intern
ml research, kernels, and the occasional peer-reviewed shitpost inference @baseten || eng @uwaterloo
San Francisco
github.com/AliesTaha
Joined October 2024
111
Following
35.4K
Followers
RepliesRepliesRepostsRepostsMediaMediaArticlesArticles

Log in or sign up for X

See what’s happening and join the conversation

Continue with phone
or
Log in with username or email
Terms·Privacy·Cookies·Accessibility·Ads Info·© 2026 X Corp.
  • Pinned
    @waterloo_intern
    ali
    Baseten
    @waterloo_intern
    Jun 26
    Article cover image
    Article
    some notes on writing the fastest video kernel in the world
    in this worklog, I explain how we retrofitted sparsity into a model at inference time, yielding the world's fastest video generation kernel. we iteratively optimize a frontier oss model by 54x the...
    35
  • @waterloo_intern
    ali
    Baseten
    @waterloo_intern
    Aug 30
    what. a. read. i’m only a little jealous
    @BrianLi23
    Brian Li
    Baseten
    @BrianLi23
    Aug 28
    Article cover image
    Article
    Agentic Kernels in Production
    TL;DR: We’ve built an agentic kernel development framework that identifies model-level optimization opportunities, generates improved kernels, and validates them in our serving stack. On our current...
    3
  • @waterloo_intern
    ali
    Baseten
    @waterloo_intern
    Aug 15
    “rumor-mills we’re following have been mentioning how the data industry is taking off in China — very much driven by American data companies selling to Chinese model labs. This could look like Chinese labs buying many of the same RL environments that are used by American frontier
    @natolambert
    Nathan Lambert
    @natolambert
    Aug 14
    GLM 5.3 notes and why we should stop being so surprised about these very strong Chinese models (most of this is talking myself through some of my denial -- yes, these models are the real deal). interconnects.ai/p/glm-53-how-c…
    7
  • @waterloo_intern
    ali
    Baseten
    @waterloo_intern
    Aug 14
    for my final post as @waterloo_intern, i'd like to ‘outroduce’ myself. it takes exactly 4 years and 8 months to make a waterloo intern. today is the last day of mine. the full life cycle, in four stages, is as follows: 1- interview hazing 2- cali or bust 3- inflection point
    Image
    47
  • @waterloo_intern
    ali
    Baseten
    @waterloo_intern
    Aug 13
    kimi paper readers in SHAMBLES after reading the deepseek paper (me, it’s me, and at least one more (henry, below)) as in, if you can get just as good (actually better of) a model with swa, how much of k3’s success can be attributed to its use of kda vs the remaining tricks, and
    @henrylhtsang
    henry tsang
    @henrylhtsang
    Aug 13
    I (very embarrasingly) only realize sliding window attention also makes kv cache bounded (i.e. indep of seq len), similar to KDA sure maybe it uses more kv cache than KDA but its not a magnitude more
    18
Advertisement
Advertisement