Log inSign up
Inferact
129 posts
Inferact profile banner
@inferact

Inferact

@inferact
Building the future of inference
San Francisco, California
inferact.ai
Joined December 2025
5
Following
6,120
Followers
RepliesRepliesRepostsRepostsMediaMedia

Log in or sign up for X

See what’s happening and join the conversation

Continue with phone
or
Log in with username or email
Terms·Privacy·Cookies·Accessibility·Ads Info·© 2026 X Corp.
  • @inferact
    Inferact
    @inferact
    20h
    Behind this blog is months of our team's work tuning vLLM on agentic workloads and validating on @SemiAnalysis_ AgentX benchmark. We find that open source models optimized for agentic workloads reach up to 130K tokens/GPU-sec, 106× cheaper than Opus 5 API pricing. vLLM is the
    @vllm_project
    vLLM
    @vllm_project
    20h
    New blog is out: vLLM x AgentX: Optimizing for Real-World Agentic Serving. Agent traffic stresses every layer of the serving stack at once. This post walks the full-stack work for optimizing vLLM on Agentic workloads, including the architecture, framework, and runtime
    Image
    00:00
    1
  • @inferact
    Inferact
    @inferact
    Sep 3
    Congratulations to @HUMAIN and @MiniMax_AI on HUMAIN-M3, a frontier Arabic model now available on HUMAIN Node 🎉. @vllm_project is running the inference under the hood, and we're looking forward to more Arabic use cases with this state-of-the-art stack.
    @HUMAIN
    HUMAIN
    @HUMAIN
    Sep 3
    HUMAIN unveils HUMAIN-M3, a frontier Arabic language model, commissioned by HUMAIN and developed by @MiniMax_AI, now available in research preview on HUMAIN Node. Read more: humain.com/news/humain-un… #HUMAIN #LEAP26 #TheEndOfLimits
    Image
  • @inferact
    Inferact
    @inferact
    Aug 28
    Our co-founder and CEO @simon_mo_ is featured in the PyTorch Conference promo, talking about where @vllm_project is headed! 🚀 Inferact and @simon_mo_ will be presenting at the conference, come find our team at San Jose, October 20-21!
    @PyTorch
    PyTorch
    @PyTorch
    Aug 28
    At PyTorch Conference North America 2026, hear directly from and connect with the people working on PyTorch, @vllm_project, @DeepSpeedAI, @raydistributed, Helion, and Safetensors, alongside experts working across the AI stack. Keynote speaker @simon_mo_ (@vllm_project,
    Image
    00:00
  • @inferact
    Inferact
    @inferact
    Aug 25
    Much of this work came out of our team at Inferact, working with @vllm_project and @SemiAnalysis_. Huge shoutout to @yifandotqiao, who led our AgentX workstream end to end, and to @esmeetu87, @LiZhewen71800, Jeff Ma, Summer Yang, Nick Hill, Woosuk Kwon, and Dao Le for months of
    @vllm_project
    vLLM
    @vllm_project
    Aug 25
    Congratulations to @SemiAnalysis_ on the release of AgentX 1.0 🎊, an open-source multi-turn agentic coding benchmark collected from ~$3M of real traces, running on 1000+ chips and ~2MW of continuously operated compute. We are excited to see @vllm_project’s competitive
    Image
    1
  • @inferact
    Inferact
    @inferact
    Aug 21
    We're excited to kick off the first vLLM Conference next week! Two days of talks and three nights of happy hours. Join an event and learn where the future of inference is heading 🚀
    @vllm_project
    vLLM
    @vllm_project
    Aug 21
    vLLM Conference is next week, and we have a packed schedule 🎊 📅. Here's the full list of events you should know: Mon: 🔷 4–6PM vLLM × Ray × Google Cloud Happy Hour: rsvp.withgoogle.com/events/ray-sum… 🔷 6–9PM vLLM × Dynamo Meetup: luma.com/r8o604o0 Tue: 🔷 11–11:30AM vLLM
    Image
    1
Advertisement
Advertisement