Your agents don't need a genius. They need a model that never becomes the bottleneck. NVIDIA Nemotron 3.5 Lightning: 30B MoE, 3B active, ~670 tokens/sec. Running on vLLM day 0, with quantized checkpoints from Red Hat AI ready now. Here's the writeup:

Article
Your agents don't need a genius. They need a model that runs at 670 tokens/sec.
Most of what an agent does all day is grunt work: tool calls, retrieval, validation, formatting, classification, summarization. You don't need a frontier reasoner for that. You need something fast,...

