Log inSign up
Machine Learning at Berkeley
249 posts
Machine Learning at Berkeley profile banner
@BerkeleyML

Machine Learning at Berkeley

@BerkeleyML
Students at UC Berkeley working on academic research, ML education, industry projects, and fostering a vibrant ML community 🧠💡
Berkeley, CA
ml.berkeley.edu
Joined December 2016
113
Following
3,726
Followers
RepliesRepliesRepostsRepostsMediaMedia

Log in or sign up for X

See what’s happening and join the conversation

Continue with phone
or
Log in with username or email
Terms·Privacy·Cookies·Accessibility·Ads Info·© 2026 X Corp.
  • @BerkeleyML
    Machine Learning at Berkeley
    @BerkeleyML
    Aug 22
    check out this paper on block sparse attention from two of our members!
    @Alex2262T
    Alexander Tian
    MatX
    @Alex2262T
    Jul 14
    Hardware yearns for block sparse attention, yet it seems largely absent from open weight LLMs. DeepSeek developed NSA, and people speculated DeepSeek v4 would integrate it, yet it was never utilized. We have a hypothesis as to why. We found that replacing dense attention with
    Image
  • @BerkeleyML
    Machine Learning at Berkeley
    @BerkeleyML
    Aug 21
    are you guilty of --dangerously-skip-permissions? check out this new intervention framework by our alum!
    @verona_teo
    Verona Teo
    @verona_teo
    Aug 6
    Deciding when to jump in and help someone—and when to hold back and let them work through it—is something humans navigate constantly. How do AI assistants handle this tradeoff? We introduce Int-Bench, a framework for evaluating interventions during problem-solving tasks.
  • @BerkeleyML
    Machine Learning at Berkeley
    @BerkeleyML
    Aug 14
    Check out this Mixture of Experts project by our member!!
    @chloewchia
    chloe chia
    @chloewchia
    Aug 13
    This week, I wanted to see if we could get the smallest possible Mixture of Experts going that takes a few hours to pretrain on a single GPU, is fast for experiments, but also somehow performs well compared to larger MoEs. For reference, we have Nanochat for transformer
    Image
  • @BerkeleyML
    Machine Learning at Berkeley
    @BerkeleyML
    Aug 11
    Check out this awesome explainer of DiffusionGemma by one of our members!
    @timg8710
    Timothy Gao
    @timg8710
    Aug 10
    This is Google’s new diffusion LLM, DiffusionGemma’s denoising canvas over time. Diffusion LLMs can generate tokens in flexible order. But in practice, do they just become autoregressive anyway? 1/ 🧵
    Image
    GIF
  • @BerkeleyML
    Machine Learning at Berkeley
    @BerkeleyML
    May 1
    our members ran a great reading group on major architecture changes in deepseek v4!! they covered: hyper connections + manifold hyper connections, KV cache + MQA/GQA intro, Deepseek Sparse attention (prerequisite to understanding the new CSA) & a walk through of CSA and HCA
    Image
    2
Advertisement
Advertisement