1. X
  2. Red Hat AI
Log inSign up
Red Hat AI
2,422 posts
Image
user avatar
Red Hat AI
@RedHat_AI
Accelerating AI innovation with open platforms and community. The future of AI is open.
ai.redhat.com
Joined May 2018
2,090
Following
11.7K
Followers
RepliesRepliesArticlesArticlesMediaMedia
  • user avatar
    Red Hat AI
    @RedHat_AI
    Jul 31
    Speculators v0.7.0 is out. DSpark is now supported, and we shipped two DSpark speculators with it: GLM-5.2 and gemma-4-31B, both with strong gains over standard autoregressive drafting. Also new: a config-file-first training CLI, Muon as the default optimizer, and ~3x faster
    Speculators v0.7.0 release summary.
    7.2K
  • user avatar
    Red Hat AI
    @RedHat_AI
    Jul 31
    Kimi K3 now in NVFP4, the Blackwell counterpart to our FP8-Block checkpoint. MoE layers quantized to 4-bit for accelerated inference on Blackwell. Quality holds, GPQA 93.5 to 91.0. huggingface.co/RedHatAI/Kimi-…
    user avatar
    Red Hat AI
    @RedHat_AI
    Jul 28
    Running Kimi K3 on Hopper? We released an FP8-Block quantized checkpoint tuned for it. FP8 is native to Hopper's tensor cores, so you get the throughput the H100/H200 are built to deliver. Day Zero support ready with @vllm_project. Weights - try it now: huggingface.co/RedHatAI/Kimi-…
    24K
  • user avatar
    Red Hat AI
    @RedHat_AI
    Jul 30
    The noisy-neighbor problem: Two teams share one model. One runs batch summarization, the other needs a live coding chatbot. The batch job saturates the GPU and the chatbot's time-to-first-token spikes. Flow control in @_llm_d_ fixes it with priority-based queuing and fairness
    Featured image for Red Hat OpenShift AI.
    Optimize GPU efficiency with OpenShift AI and llm-d flow-control | Red Hat Developer
    From developers.redhat.com
    2.1K
  • user avatar
    Red Hat AI
    @RedHat_AI
    Jul 29
    Speculators and @vllm_project now support parallel drafting: P-EAGLE, DFlash, DSpark predict a whole block of draft tokens in one pass. No sequential penalty, still lossless. Blog: vllm.ai/blog/2026-07-2… DFlash on Qwen3-30B-A3B for coding, compared to EAGLE-3 and no spec:
    Parallel drafting algorithms, such as P-EAGLE, DFlash and DSpark, provide significant performance gains when compared to autoregressive drafting algorithms such as EAGLE-3. Speculator models mentioned above can be found in the Speculators Collection at the RedHatAI HuggingFace Hub.
    2.9K
  • user avatar
    Red Hat AI
    @RedHat_AI
    Jul 28
    Kimi K3 open weights landed yesterday. By that afternoon, Red Hat AI Inference was serving it on one 8x B300 node. 2.8T params. Day-0 preview images exist so you can experiment the moment weights drop. 👇
    Article cover image
    Article
    Run Kimi K3 on Day 0 with Red Hat AI. Here's the Exact Command.
    Kimi K3 dropped open weights on July 27th. By that afternoon, we had it serving an OpenAI-compatible API on a single 8×B300 node using Red Hat AI Inference. This is the largest open-weight model ever...
    3.9K
  • See @RedHat_AI's full profile

    Sign up
    Log in

Log in or sign up for X

See what’s happening and join the conversation

Continue with phone
or
Log in with username or email
Terms·Privacy·Cookies·Accessibility·Ads Info·© 2026 X Corp.
Advertisement
Advertisement