1. X
  2. vLLM
Log inSign up
vLLM
1,232 posts
vLLM profile banner
user avatar

vLLM

@vllm_project
A high-throughput and memory-efficient inference and serving engine for LLMs. Join slack.vllm.ai to discuss together with the community!
vllm.ai
Joined March 2024
36
Following
46.6K
Followers
RepliesRepliesArticlesArticlesMediaMedia
  • Pinned
    user avatar
    vLLM
    @vllm_project
    Jul 9
    Announcing the first-ever vLLM Conference β€” hosted by @inferact at Ray Summit, Aug 24–26 in San Francisco πŸŽ‰πŸŒ‰ This is where we'll get into the work pushing open, high-performance inference forward, such as: πŸ—ΊοΈ Where the vLLM roadmap is headed ⚑ Getting the most out of
    Image
    First vLLM Conference at Ray Summit
    From vllm.ai
  • user avatar
    vLLM
    @vllm_project
    5h
    Good to see @RedNote back with a new open model, and a much bigger one built for long-horizon agent work. dots3-note preview is a 280B MoE with 16B active at 512K context, it reads images, audio, and video, and it is trained to explore unfamiliar environments and update its own
    Image
    Image
    Image
    user avatar
    dots studio
    @dotsstudioai
    7h
    Introducing dots3-note preview β€” a small but mighty step toward long-horizon agency in real life. πŸ”Ή 280B MoE with 16B active parameters, a 512K context window, and multimodal understanding across text, vision, and audio πŸ”Ή Introduces TEMPO, a new RL approach for long-horizon
  • user avatar
    vLLM
    @vllm_project
    Aug 12
    πŸŽ‰ Congrats to @Alibaba_Qwen on Qwen3.8-2.4T-A95B, one of the largest open-weight models released to date. 2.4T params, 95B active, 512 experts. Day-0 support in vLLM, verified on @nvidia and @AMD hardware. A ready-made 4-bit checkpoint per vendor, both out of the box:
    Image
    Day 0 Support for Qwen3.8-2.4T-A95B on vLLM
    From vllm.ai
  • user avatar
    vLLM
    @vllm_project
    Aug 12
    vLLM's model loader and KV connector both have an @Azure Blob path now: weights in, KV out. @Microsoft and @NVIDIAAI shipped a recipe for each. Nothing serves until weights land in HBM. Dynamo ModelExpress plugs into the loader, up to 7.3x faster than the default on H100/A100.
    Image
  • user avatar
    vLLM
    @vllm_project
    Aug 11
    We're co-hosting a @vllm_project x @nvidia Dynamo Meetup during vLLM Conference week in San Francisco ⚑ Tech talks on serving LLMs efficiently at scale + food, drinks, and the inference community all in one room πŸ“… Aug 24, 6–9pm PT Space is limited β€” RSVP:
    Image
    vLLM x NVIDIA Dynamo Meetup Β· Luma
    From luma.com

Log in or sign up for X

See what’s happening and join the conversation

Continue with phone
or
Log in with username or email
TermsΒ·PrivacyΒ·CookiesΒ·AccessibilityΒ·Ads InfoΒ·Β© 2026 X Corp.
Advertisement
Advertisement