Log inSign up
Ettore Di Giacinto
3,114 posts
Ettore Di Giacinto profile banner
@mudler_it

Ettore Di Giacinto

@mudler_it
dad, creator of LocalAI(localai.io) and Kairos (kairos.io) , ex @SUSE/@Rancher, ex-Gentoo Dev. vllm.cpp and APEX quants
Italy
github.com/mudler
Joined January 2016
268
Following
4,860
Followers
RepliesRepliesRepostsRepostsMediaMedia

Log in or sign up for X

See what’s happening and join the conversation

Continue with phone
or
Log in with username or email
Terms·Privacy·Cookies·Accessibility·Ads Info·© 2026 X Corp.
  • Pinned
    @mudler_it
    Ettore Di Giacinto
    @mudler_it
    Aug 13
    a week ago I shared what I've been working on, vllm.cpp, vLLM's serving stack in C++ with no python in it. I got really a lot of feedback, and it seems I'm not the only one who wanted this to exist. We reached 800+ commits, 280 stars in just so little time. but where is it
    Image
    9
  • @mudler_it
    Ettore Di Giacinto
    @mudler_it
    1h
    vllm.cpp is now getting native exl3 support.
    @MiaAI_lab
    Mia
    @MiaAI_lab
    16h
    Rejoice RTX 3090/4090/5090 owners💫 Qwen3.8-27B EXL3 is here. This is an *experimental* release so there might be some bugs along the way! As far as I know, this is the only way to serve Qwen3.8-27B with 200k+ context while using dflash2 on a 24GB VRAM, without sacrificing
    Image
  • @mudler_it
    Ettore Di Giacinto
    @mudler_it
    Aug 30
    Hey @github . Can you provide some clarity on your ToS? It's so _silly_ that in the AI era we can't have separate accounts for operating with agents. This is so dumb that I can't take more of it. I don't complain about your uptime status, but this is very nonsense. I got
    4
  • @mudler_it
    Ettore Di Giacinto
    @mudler_it
    Aug 28
    Thanks @AMD 🫶 They have sent a box over... so this means you are gonna get first class @AMD support in vllm.cpp, @LocalAI_API and all our projects. It was very painful before, as we had to work only with logs and reported issues, expects things to improve a lot of you have one
    Image
    12
  • @mudler_it
    Ettore Di Giacinto
    @mudler_it
    Aug 26
    Qwen Flash next is here, 180b with 6b active. This when quantized correctly is gonna be a real workforce for DGX and alikes.
    Image
    Qwen/Qwen3.8-Flash-Next · Hugging Face
    From huggingface.co
    6
Advertisement
Advertisement