Log inSign up
llm-d
248 posts
llm-d profile banner
@_llm_d_

llm-d

@_llm_d_
llm-d: a Kubernetes-native high-performance distributed LLM inference framework
llm-d.ai
Joined May 2025
2
Following
906
Followers
RepliesRepliesRepostsRepostsMediaMedia

Log in or sign up for X

See what’s happening and join the conversation

Continue with phone
or
Log in with username or email
Terms·Privacy·Cookies·Accessibility·Ads Info·© 2026 X Corp.
  • Pinned
    @_llm_d_
    llm-d
    @_llm_d_
    Sep 8
    Great open-source communities push enterprise hardware further. 🤝 New work from IBM Research & Red Hat on llm-d: • 753B open model on 544 NVIDIA H100 GPUs • 5–10x lower cost per token vs commercial APIs • Serves 1000s of concurrent agents Blog:
    Image
    How llm-d makes the most of the hardware you already have
    From research.ibm.com
  • @_llm_d_
    llm-d
    @_llm_d_
    Aug 20
    AI is entering the era of large-scale, distributed, agentic systems. Applications coordinate multiple models, tools, and services, process millions of requests, and demand enormous amounts of computer capacity.
    Image
    redhat.com
    Scaling agentic AI: How llm-d enables infrastructure sovereignty
    Scale agentic AI with llm-d to achieve full infrastructure sovereignty across diverse hardware.
  • @_llm_d_
    llm-d
    @_llm_d_
    Aug 18
    Sticky Until Saturated: Token-Aware Routing in llm-d The llm-d router's default configuration is built on token-aware routing: keep each request on the cache-warm endpoint until a calibrated token-load limit is exceeded, then route by load alone ...
    Image
    Sticky Until Saturated: Token-Aware Routing in llm-d | llm-d
    From llm-d.ai
  • @_llm_d_
    llm-d
    @_llm_d_
    Aug 18
    Benchmarking disaggregated VLM serving of Kimi-VL-A3B-Instruct with 4 Intel Arc Pro B60 vision encoder and 1 NVIDIA H200 language model worker, delivering 2.4x-2.8x higher throughput and 69%-80% lower TTFT ... llm-d.ai/blog/scaling-v…
    Image
  • @_llm_d_
    llm-d
    @_llm_d_
    Jul 27
    Check out this great summary of the recent @vllm_project office hours.
    @RedHat_AI
    Red Hat AI
    @RedHat_AI
    Jul 23
    Article cover image
    Article
    Distributed Inference and Wide Expert Parallelism with llm-d: vLLM Office Hours #53 Recap
    At vLLM Office Hours #53, Robert Shaw, vLLM core maintainer and llm-d lead maintainer at Red Hat AI, walked through what llm-d is, how it orchestrates distributed LLM inference on Kubernetes, and why...
Advertisement
Advertisement