1. X
  2. Inferact
Log inSign up
Inferact
110 posts
Image
user avatar
Inferact
@inferact
Building the future of inference
San Francisco, California
inferact.ai
Joined December 2025
5
Following
5,327
Followers
RepliesRepliesMediaMedia

Log in or sign up for X

See what’s happening and join the conversation

Continue with phone
or
Log in with username or email
Terms·Privacy·Cookies·Accessibility·Ads Info·© 2026 X Corp.
  • Pinned
    user avatar
    Inferact
    @inferact
    Jul 9
    Proud to host the first vLLM Conference at Ray Summit, Aug 24–26 in SF 🚀🎉 We're bringing the community together to talk about what's next: 🏗️ Building and scaling AI in production 🌐 Deploying vLLM to run inference across your own cloud and hardware 🤖 What open source means
    user avatar
    vLLM
    @vllm_project
    Jul 9
    Announcing the first-ever vLLM Conference — hosted by @inferact at Ray Summit, Aug 24–26 in San Francisco 🎉🌉 This is where we'll get into the work pushing open, high-performance inference forward, such as: 🗺️ Where the vLLM roadmap is headed ⚡ Getting the most out of
    6.5K
  • user avatar
    Inferact
    @inferact
    Jul 29
    Highlighting the amazing work of our team. On a low entropy workload, we reached 464 tok/s decode on Kimi K3 + DSpark at batch size 1. Shoutout to @zixi_qi_ and the Inferact team for this work: pushing @vllm_project performance to the speed of light 🚀
    user avatar
    vLLM
    @vllm_project
    Jul 29
    vLLM hit new peak bs=1 decode on Kimi-K3: 464 tok/s 🚀 Under a low-entropy reasoning workload, Kimi-K3 + DSpark on vLLM reaches 464 tok/s on batch size 1 with 4×4 GB300. This benchmark is fully reproducible with public image: vllm/vllm-openai:kimi-k3 and @inferact's DSpark
    Image
    6.9K
  • user avatar
    Inferact
    @inferact
    Jul 27
    Kimi K3 is one of the most powerful open-weight models ever released: 2.8T params, 1M context, native vision. Getting it to serve well on day 0 took real engineering. The @inferact team led the vLLM integration: KDA-aware caching, MXFP4 MoE kernels, and a DSpark
    user avatar
    vLLM
    @vllm_project
    Jul 27
    With Kimi K3 Day-0 on vLLM: Open Frontier Intelligence for Everyone 🚀 At 2.8 trillion parameters, Moonshot AI's Kimi K3 is one of the most powerful open-weight models ever released. Starting today, you can serve it on vLLM the moment the weights are public. What K3 brings: 🧠
    Image
    00:00
    2.5K
  • user avatar
    Inferact
    @inferact
    Jul 26
    Inferact is proud to co-sign. We were founded to grow @vllm_project into the world’s leading open-source inference engine, making open-weight AI faster, cheaper, and accessible across hardware and clouds so intelligence can be on tap for everyone.
    Image
    Image
    Image
    Image
    user avatar
    Jensen Huang
    NVIDIA
    @JensenHuang
    Jul 24
    For my first post, I’m sharing a letter @nvidia signed on why open models matter. AI will transform every industry, power every company, and be built by every country. Open models strengthen safety and cybersecurity, accelerate innovation and diffusion, and enable sovereignty.
    31K
  • user avatar
    Inferact
    @inferact
    Jul 15
    Day 0 and vLLM already runs Inkling at full feature parity. ⚡ We worked with @thinkymachines to bring their 1T-parameter multimodal model to vLLM on launch day — text, image, and audio in, with up to 1M context. Performance: up to 380 tok/s/user with MTP on 4× GB200s.
    user avatar
    Thinking Machines
    @thinkymachines
    Jul 15
    Today, we are introducing Inkling. Inkling reasons efficiently across text, image, and audio modalities. We are making the full weights available. thinkingmachines.ai/news/introduci… Available today for fine-tuning on Tinker. Play with it in the Inkling Playground. 🧵
    4.2K
  • See @inferact's full profile

    Sign up
    Log in
Advertisement
Advertisement