Log inSign up
vLLM
1,273 posts
vLLM profile banner
@vllm_project

vLLM

@vllm_project
A high-throughput and memory-efficient inference and serving engine for LLMs. Join slack.vllm.ai to discuss together with the community!
vllm.ai
Joined March 2024
36
Following
48.4K
Followers
RepliesRepliesRepostsRepostsMediaMediaArticlesArticles

Log in or sign up for X

See what’s happening and join the conversation

Continue with phone
or
Log in with username or email
Terms·Privacy·Cookies·Accessibility·Ads Info·© 2026 X Corp.
  • @vllm_project
    vLLM
    @vllm_project
    13h
    🎉 Congrats to @AntLingAGI on open-sourcing Ling-3.0-flash-Fin and FinFIRST! Open models and expert-built benchmarks are a great step toward more accessible and verifiable financial AI. Looking forward to seeing these workflows run at scale with vLLM. 🚀
    @AntLingAGI
    Ant Ling
    @AntLingAGI
    22h
    We’re open-sourcing Ling-3.0-flash-Fin, a finance-enhanced model for real-world workflows, and FinFIRST, an expert-built benchmark for financial search agents. Two open releases, one goal: making financial AI more accessible and verifiable.
    Image
    1
  • @vllm_project
    vLLM
    @vllm_project
    13h
    Love seeing @lightseekorg use TorchSpec + vLLM to train three different K3 draft models, with the recipes open sourced too. Exactly the kind of training-serving collaboration we love to see 🚀
    @lightseekorg
    LightSeek Foundation
    @lightseekorg
    22h
    Releasing the @Kimi_Moonshot K3 Draft Collection — 3 draft models (EAGLE-3, DFlash2, DSpark) trained with TorchSpec and @vllm_project on @NVIDIAAI GB200. 🚀 We also shared the data recipes. More on the blog → lightseek.org/blog/kimi-k3-d…
    Image
    1
  • @vllm_project
    vLLM
    @vllm_project
    17h
    Thank you to the @SemiAnalysis_ team for the shoutout and for the collaboration on AgentX 🙏 Benchmarks are only useful when they measure the workloads people actually run, and AgentX measures the real thing: multi-turn, long-context agent traffic. vLLM is the engine for
    @SemiAnalysis_
    SemiAnalysis
    @SemiAnalysis_
    18h
    Shoutout to the cracked team at @vllm_project that implemented recent agentic workload optimizations. (1/5)🧵
    Image
    1
  • @vllm_project
    vLLM
    @vllm_project
    23h
    K2-Horizon has day-0 support in vLLM, and IFM released intermediate checkpoints, detailed data-construction recipes, the training code, and fine-grained logs alongside the weights. 📷 512K context and Apache-2.0 from 3.7B up, with reasoning and tool calling in the checkpoints
    @IFM_AI
    Institute of Foundation Models
    @IFM_AI
    Sep 3
    Introducing K2 Horizon: a connected fleet of six foundation models ranging from 0.9 billion to 375 billion parameters. - Frontier performance: Across coding and agentic tasks, K2 Horizon delivers top-tier performance in every size class—with the 0.9B, 3.7B and 7B models setting
    Image
    3
  • @vllm_project
    vLLM
    @vllm_project
    Sep 1
    🎬Video generation faster than playback! 🚀MiniMax H3 on vLLM-Omni + FastVideo's FastH3: a complete 10.1s MP4 - video AND synchronized audio - rendered in 8.7s!⚡️ Thanks to @MiniMax_AI for the great Minimax H3 release, the FastVideo team @haoailab for open-sourcing FastH3 and
    Image
    22
Advertisement
Advertisement