1. X
  2. Red Hat AI
Log inSign up
Red Hat AI
2,434 posts
Image
user avatar
Red Hat AI
@RedHat_AI
Accelerating AI innovation with open platforms and community. The future of AI is open.
ai.redhat.com
Joined May 2018
2,090
Following
11.8K
Followers
RepliesRepliesArticlesArticlesMediaMedia
  • user avatar
    Red Hat AI
    @RedHat_AI
    4h
    We're back with @vllm_project office hours this Thursday, August 6! We'll share: ๐Ÿš€ What's new in vLLM 0.26 ๐Ÿ† A winning project from a recent vLLM hackathon โšก A deep dive into Mooncake + vLLM/llm-d, the infra behind 80% of Kimi's traffic, and how its distributed KV cache pool
  • user avatar
    Red Hat AI
    @RedHat_AI
    5h
    Unmonitored AI training jobs cost time and compute. Progress tracking in Red Hat OpenShift AI brings real-time visibility to every training step, helping teams protect investments and optimize hybrid cloud performance. See how to maximize every GPU hour:
    Image
    redhat.com
    Make every GPU-hour count: Progress tracking in Red Hat OpenShift AI
    Save GPU hours with Red Hat OpenShift AI's progress tracking for ML training jobs
  • user avatar
    Red Hat AI
    @RedHat_AI
    Aug 4
    A coding agent's turn is mostly reading. Across 219 real Claude Code sessions from @SemiAnalysis_, the median request sends 195K tokens in and 317 out, and 96% of turns reuse at least 90% of their prior context. 98.6% of all tokens served are input. That's the workload @_llm_d_
    Highest measured concurrency that sustains at least 30 average tok/s/user. MTP meets the target through c128 and serves 5,647 tok/s; CPU offloading alone meets it only at c16 and serves 561 tok/s. That is at least 8x the concurrency and 10.1x the aggregate output. The target is an average output-rate floor, not a TTFT or tail-latency bound.
  • user avatar
    Red Hat AI
    @RedHat_AI
    Aug 4
    Most RAG configs are a guess: grab a chunk size from a tutorial, set top-k high to be safe, ship it. On a small model that backfires, it pulls the wrong number out of a pile of lookalikes. AutoRAG sweeps the settings and scores retrieval with no LLM calls, so testing 144 configs
    Image
    00:00
  • user avatar
    Red Hat AI
    @RedHat_AI
    Aug 3
    Same model, same two H100s. The only thing that changes is precision. Llama 3 70B at FP8 instead of FP16: throughput goes from 158 to 474 tokens/sec, and time-to-first-token under load drops from ~30s to under 5s. @cedricclyburn guide walks through how quantization gets you
    user avatar
    Cedric Clyburn
    @cedricclyburn
    Aug 3
    Article cover image
    Article
    The LLM Quantization Guide: How it works & the benefits
    Model sizes have roughly doubled every year, and GPU memory hasn't come close to keeping up. That gap is why most models in production are quantized, and I'll explain here: What quantization actually...

Log in or sign up for X

See whatโ€™s happening and join the conversation

Continue with phone
or
Log in with username or email
TermsยทPrivacyยทCookiesยทAccessibilityยทAds Infoยทยฉ 2026 X Corp.
Advertisement
Advertisement