1. X
  2. Ettore Di Giacinto
Log inSign up
Ettore Di Giacinto
3,012 posts
Ettore Di Giacinto profile banner
user avatar

Ettore Di Giacinto

@mudler_it
dad, creator of LocalAI(localai.io) and Kairos (kairos.io) , ex @SUSE/@Rancher, ex-Gentoo Dev.
Italy
github.com/mudler
Joined January 2016
265
Following
3,672
Followers
RepliesRepliesMediaMedia
  • Pinned
    user avatar
    Ettore Di Giacinto
    @mudler_it
    Aug 7
    vllm.cpp runs @MiniMax_AI 's MiniMax-H3 now. 33.1B, video and audio out of a single model, and I'm still a bit stunned that it works we reimplemented both VAEs from the checkpoint's remote python, so there's no torch anywhere in the process. And you drive the whole thing over
    Image
    00:00
  • user avatar
    Ettore Di Giacinto
    @mudler_it
    Aug 10
    vllm.cpp on GTX5090 Ti @jichiep is unstoppable!
    user avatar
    Richard Palethorpe
    @jichiep
    Aug 10
    Over the weekend I did some more optimization for my card in vllm.cpp and pushed throughput beyond vLLM for the first time. This is all done by following the optimization protocol layed out by @mudler_it. It's relatively easy to add support for a card and optimize it, but if you
    Image
  • user avatar
    Ettore Di Giacinto
    @mudler_it
    Aug 10
    @tenstorrent is getting as well into vllm.cpp! Luca is really underselling this, because in just a couple of hours we got coherent text! From there it can only get better. Thanks to the community vllm.cpp is moving very quickly
    user avatar
    Luca Barbato
    @lu_zero_
    Aug 10
    This weekend I started adding support for @tenstorrent Blackhole on vllm.cpp, it is fairly slow but is progressing github.com/mudler/vllm.cp… Next weekend I'll try to add more.
  • user avatar
    Ettore Di Giacinto
    @mudler_it
    Aug 10
    RDNA4 is getting shape in vllm.cpp!
    user avatar
    Mike
    @anothervariable
    Aug 10
    Lots to chew on here and it's painfully slow. But it is working on vllm.cpp/rocm Gemma 4 26B MoE :). What you see in the screenshot is vllm.cpp with rocm and ported cuda kernels into rocm/hip RDNA4 serving Gemma 4 26B MoE. Prefill and decoding and responding to Hermes Agent
    Image
  • user avatar
    Ettore Di Giacinto
    @mudler_it
    Aug 8
    this is what I hoped would happen, and not this fast two days ago I shipped an AMD backend in vllm.cpp that had never been compiled by anyone, including me, because I have no AMD card. it went in as a skeleton to get things started, so everyone could build on top. now there's
    user avatar
    Mike
    @anothervariable
    Aug 8
    Replying to @mudler_it and @AMD
    Two PRs later and we have Gemma 4 26B working in vllm.cpp and rocm using 2xR9700 GPUs 😁. Yes Grok is driving 😁.
    Image

Log in or sign up for X

See what’s happening and join the conversation

Continue with phone
or
Log in with username or email
Terms·Privacy·Cookies·Accessibility·Ads Info·© 2026 X Corp.
Advertisement
Advertisement