MAX is Modular's inference stack for running any model on any chip: one codebase, one model definition, one kernel library. High-performance AI serving and modeling for any hardware.
In this technical deep dive from ModCon 2026, Senior AI Product Manager @ehsanmok and Head of
People ask what our secret is when they see our performance numbers.
Brendan Hansknecht, AI Performance Engineering Manager, on why performance is a full-stack problem, and why gluing together someone else's kernels, tokenizers, and schedulers doesn't work in production:
Support for new AI hardware usually takes a big team, many repos, and a year.
Two and a half engineers got a frontier open model serving on @Qualcomm Cloud AI 100 in under six months using the Modular stack, followed by GPT-2 on Qualcomm Dragonfly™ AI 200 in a week.
Inside the
GLM-5.3 open weights are now public, and Modular Cloud has Day Zero support.
GLM-5.3 is the most capable open-weights model for coding, with a 50% improvement over GLM-5.2 on @Zai_org's Code Bench.
Try it today on Modular Cloud: console.modular.com/?utm_source=x&…