Log inSign up
Simon Mo
Inferact
480 posts
@simon_mo_

Simon Mo

Inferact
@simon_mo_
building @inferact for @vllm_project
Joined July 2018
368
Following
4,193
Followers
RepliesRepliesRepostsRepostsMediaMediaArticlesArticles

Log in or sign up for X

See what’s happening and join the conversation

Continue with phone
or
Log in with username or email
Terms·Privacy·Cookies·Accessibility·Ads Info·© 2026 X Corp.
  • Pinned
    @simon_mo_
    Simon Mo
    Inferact
    @simon_mo_
    Jan 22
    vLLM has grown to 2000+ contributors scale with a diverse community of model, hardwares, and applications. I see @vllm_project on the path of becoming the world's inference engine and @inferact to accelerate AI progress. We cannot be more excited about the road ahead.
    @woosuk_k
    Woosuk Kwon
    Inferact
    @woosuk_k
    Jan 22
    Today, we're proud to announce @inferact, a startup founded by creators and core maintainers of @vllm_project, the most popular open-source LLM inference engine. Our mission is to grow vLLM as the world's AI inference engine and accelerate AI progress by making inference cheaper
    Image
    12
  • @simon_mo_
    Simon Mo
    Inferact
    @simon_mo_
    22h
    🫡 vLLM is your engine of choice for agentic workload. The pareto curve is easily understood; but very hard to optimize against. The @inferact team did amazing work here, please check it out (and join us!)
    @vllm_project
    vLLM
    @vllm_project
    23h
    New blog is out: vLLM x AgentX: Optimizing for Real-World Agentic Serving. Agent traffic stresses every layer of the serving stack at once. This post walks the full-stack work for optimizing vLLM on Agentic workloads, including the architecture, framework, and runtime
    Image
    00:00
  • @simon_mo_
    Simon Mo
    Inferact
    @simon_mo_
    Sep 3
    IYKYK
    Image
    @SemiAnalysis_
    SemiAnalysis
    @SemiAnalysis_
    Sep 3
    Image
    Shoutout to the cracked team at @vllm_project that implemented recent agentic workload optimizations. (1/5)🧵
    3
  • @simon_mo_
    Simon Mo
    Inferact
    @simon_mo_
    Aug 30
    Production quality is a bar, as a community, we should never compromise. Great collaboration between @vllm_project and @FireworksAI_HQ investigating this! 🫡
    @lqiao
    Lin Qiao
    Fireworks
    @lqiao
    Aug 29
    GLM-5.3-Flash is live on Fireworks. Day 2. Why not Day 0? Because being first isn't the goal. Being correct is. We unplugged peculiar behaviors of over-thinking from the initial tests, and shared all fixes back to open source. Quality trumps hype. Let's hold a high quality
    2
  • @simon_mo_
    Simon Mo
    Inferact
    @simon_mo_
    Aug 26
    Our kernel god: K3 mogs Codex Sol
    @gaunernst
    Thien Tran
    Inferact
    @gaunernst
    Aug 26
    Recently I tried using K3 for Triton kernels and I was surprised by how good it was. The generated kernels are definitely very different from how Codex would have done it. Not a proper experiment but I asked Codex (Sol medium) and K3 (high) to work a the same problem. K3 came out
    Image
    1
Advertisement
Advertisement