This Asari approach works, imagine the improvement in inference and the reduction of data center serving costs when people adopt this! The power of agents that are incredibly smart.
Our self-improving agents optimized the full @vllm_project inference stack, with up to 16% more throughput and interactivity for @deepseek_ai v4 Pro and @Zai_org GLM 5.2 on B200s (no MTP).
Every change was verified and our agents got better and faster at it with each iteration.




