Log inSign up
mobicham
643 posts
mobicham profile banner
@mobicham

mobicham

@mobicham
I like to shrink dem models 🤏 ML/AI Perf @dropbox Prev. Co-Founder & Principal Scientist @mobius_labs (acquired by @dropbox) PhD @inria
Berlin, Germany
linkedin.com/in/hicham-badr…
Joined November 2023
121
Following
725
Followers
RepliesRepliesRepostsRepostsMediaMedia

Log in or sign up for X

See what’s happening and join the conversation

Continue with phone
or
Log in with username or email
Terms·Privacy·Cookies·Accessibility·Ads Info·© 2026 X Corp.
  • @mobicham
    mobicham
    @mobicham
    Aug 20
    Gemma 4 with image inputs is very slow with the Triton attention backend because it doesn't prune the sliding window. Small fix, 3-4x faster E2E inference 🫡
    Image
    [Kernel][Gemma4] Prune Triton sliding-window tiles for multimodal prefixes by mobicham · Pull...
    From github.com
  • @mobicham
    mobicham
    @mobicham
    Jul 29
    When you ask Opus 5 a simple question 💀
    Image
    2
  • @mobicham
    mobicham
    @mobicham
    Jul 16
    The latest vLLM releases seem to be broken: randomly freezing and getting stuck at Avg prompt throughput: 0 tokens/s, and it doesn't seem related to the new V2 runner 🤔. Anyone experiencing a similar issue?
    4
  • @mobicham
    mobicham
    @mobicham
    Jul 15
    HQQ spotted 👀
    @hu_yifei
    Yifei Hu
    Reducto
    @hu_yifei
    Jul 14
    the model is smaller than the dspark drafter now. Awesome work!
    Image
  • @mobicham
    mobicham
    @mobicham
    Jun 29
    flash attention seems to be broken on rocm. Qwen3.5 training goes nowhere with fa2. With the exact same code, switching to sdpa solves the issue. Anyone familiar with this?
Advertisement
Advertisement