.@hippocraticai's health agents call tens of thousands of patients a day. Each conversational turn has to finish in about 800ms or the call stops feeling human.
That constraint is why they came to us.
Today's AI software is fragmented. Every accelerator has its own compiler, kernel libraries, runtime, and development path. Developers absorb the cost of this fragmentation.
In his ModCon 2026 tech talk, Abdul Dakkak, Chief Scientist at Modular, presents our alternative: a
MAX is Modular's inference stack for running any model on any chip: one codebase, one model definition, one kernel library. High-performance AI serving and modeling for any hardware.
In this technical deep dive from ModCon 2026, Senior AI Product Manager @ehsanmok and Head of
People ask what our secret is when they see our performance numbers.
Brendan Hansknecht, AI Performance Engineering Manager, on why performance is a full-stack problem, and why gluing together someone else's kernels, tokenizers, and schedulers doesn't work in production: