Frontier models require frontier speculators. Our team moved mountains to achieve 460 tok/s with our custom DFlash model.
Try it day 0 at modal.com/endpoints.
❤️s/RTs are randomized and differentially private.
- sitting in a rooftop restaurant in Shibuya, contemplating the time I trekked to @berniemoreno’s Mercedes Benz dealership on the Cleveland automile to convince him to reorient his Cleveland tech campaign “Blockland” to be more AI focused what a strange timelineI was just briefed by the White House on what to expect this evening. I would encourage every American to tune in tonight to the President’s speech. This may be the most important Oval Office address since the Cuban Missile Crisis. The time for complacency with China is over.
- Modal Auto Endpoints provide state-of-the-art open source inference perf with a click. Learn how we developed our low latency inference playbook with @DecagonAI, delivering responses 60ms faster than the best proprietary provider. modal.com/blog/achieve-s…
- Engram is one of the more sophisticated teams I’ve had the pleasure of working with at Modal exciting launch, looking forward to assisting with even weirder deployments in the future!Modal's been super important for our velocity over the last 6 months - Training on each user's context means scaling out to thousands of GPUs in quick bursts. Modal allowed us to do this from day zero, before we could keep a large committed cluster hot - Our research team








