Log inSign up
Woosuk Kwon
Inferact
384 posts
@woosuk_k

Woosuk Kwon

Inferact
@woosuk_k
@inferact | @vllm_project | prev: PhD @Berkeley_EECS
woosuk.me
Joined April 2023
824
Following
8,796
Followers
RepliesRepliesRepostsRepostsMediaMedia

Log in or sign up for X

See what’s happening and join the conversation

Continue with phone
or
Log in with username or email
Terms·Privacy·Cookies·Accessibility·Ads Info·© 2026 X Corp.
  • Pinned
    @woosuk_k
    Woosuk Kwon
    Inferact
    @woosuk_k
    Jan 22
    Today, we're proud to announce @inferact, a startup founded by creators and core maintainers of @vllm_project, the most popular open-source LLM inference engine. Our mission is to grow vLLM as the world's AI inference engine and accelerate AI progress by making inference cheaper
    Image
    183
  • @woosuk_k
    Woosuk Kwon
    Inferact
    @woosuk_k
    Sep 8
    TPU is an interesting chip with enormous potential. Shoutout to @googlecloud for all great work and support!
    @dylan522p
    Dylan Patel
    SemiAnalysis
    @dylan522p
    Sep 7
    We are excited to bring the first open benchmarking of Google's TPUs to the world Running every day, on many models + scenarios $/token is better than B200 and B300 Huge shout-out to Google @inferact and the InferenceX team at SemiAnalysis to this effort that's taken many months
    4
  • @woosuk_k
    Woosuk Kwon
    Inferact
    @woosuk_k
    Aug 11
    We are hiring! Come join us and advance the frontier of AI inference!
    @SemiAnalysis_
    SemiAnalysis
    @SemiAnalysis_
    Aug 10
    The @vllm_project maintainers at @inferact 🚀 are some of the most cracked engineers in the world. They’re building one of the inference engines that powers much of the world’s intelligence—and doing so with remarkable dedication, kindness, and hard work.
    Image
    5
  • @woosuk_k
    Woosuk Kwon
    Inferact
    @woosuk_k
    Aug 4
    See you at the conference!
    @vllm_project
    vLLM
    @vllm_project
    Aug 4
    The vLLM Conference is coming up in 3 weeks! 🎉 Come learn about the current state and future of AI inference, Aug 24–26 in San Francisco 🌉, hosted by @inferact at @anyscalecompute Ray Summit. We'll have speakers from Inferact, NVIDIA, AMD, Google TPU, Anyscale, PyTorch, Meta,
    Image
    2
  • @woosuk_k
    Woosuk Kwon
    Inferact
    @woosuk_k
    Jul 30
    Great to see the model small enough to run on a single (!) B300. It might be the best model in its size class! Have fun 🫡
    @vllm_project
    vLLM
    @vllm_project
    Jul 30
    🎉 Congrats to @thinkymachines on Inkling-Small-- live with Day 0 vLLM support! 276B total parameters with 12B active, native text, image, and audio input, and a 1M-token context window. Open weights, built for agentic and tool-use systems, coding assistants, and RAG. Running
    Image
    2
Advertisement
Advertisement