1. X
  2. Red Hat AI
Log inSign up
Red Hat AI
2,462 posts
Red Hat AI profile banner
user avatar

Red Hat AI

@RedHat_AI
Accelerating AI innovation with open platforms and community. The future of AI is open.
ai.redhat.com
Joined May 2018
2,090
Following
12K
Followers
RepliesRepliesArticlesArticlesMediaMedia
  • user avatar
    Red Hat AI
    @RedHat_AI
    19h
    Your agents don't need a genius. They need a model that never becomes the bottleneck. NVIDIA Nemotron 3.5 Lightning: 30B MoE, 3B active, ~670 tokens/sec. Running on vLLM day 0, with quantized checkpoints from Red Hat AI ready now. Here's the writeup:
    Article cover image
    Article
    Your agents don't need a genius. They need a model that runs at 670 tokens/sec.
    Most of what an agent does all day is grunt work: tool calls, retrieval, validation, formatting, classification, summarization. You don't need a frontier reasoner for that. You need something fast,...
  • user avatar
    Red Hat AI
    @RedHat_AI
    Aug 13
    How does a single-GPU RAG chatbot become a multi-workload inference platform? Grace from Red Hat AI walks through the whole stack: AI Gateway, vLLM, and llm-d. Quantization, speculative decoding, cache-aware routing, and prefill/decode disaggregation, one step at a time. In
    Image
    00:00
  • user avatar
    Red Hat AI
    @RedHat_AI
    Aug 13
    ~4x faster Kimi-K3 decoding. Our new DSpark speculator takes single-stream interactivity from ~110 to ~435 tok/s/user on math reasoning, delivers ~3.5x the output throughput at matched interactivity under load, and thanks to sliding window attention (2048-token window across all
    Image
  • user avatar
    Red Hat AI
    @RedHat_AI
    Aug 12
    Meta's Muse-Glimmer, running with Red Hat AI Inference (early preview). Copy the manifest, oc apply, and you're serving the multimodal model on @vllm_project. FP8, INT4, and NVFP4 checkpoints ready on Hugging Face for production. Here's the full setup in under 2 minutes by
    Image
    00:00
  • user avatar
    Red Hat AI
    @RedHat_AI
    Aug 12
    Two more Muse-Glimmer 30B checkpoints, quantized with LLM Compressor for @vllm_project: → W4A16: 4-bit weights, group size 128. huggingface.co/RedHatAI/Muse-… → NVFP4: 4-bit weights and activations, tuned for Blackwell. huggingface.co/RedHatAI/Muse-… Vision tower kept in full precision on
    Image
    RedHatAI/Muse-Glimmer-30B-W4A16 · Hugging Face
    From huggingface.co

Log in or sign up for X

See what’s happening and join the conversation

Continue with phone
or
Log in with username or email
Terms·Privacy·Cookies·Accessibility·Ads Info·© 2026 X Corp.
Advertisement
Advertisement