Frontier models are simply too expensive and slow for the majority of use cases, so we see models like SWE-2 becoming the daily driver for most.
Training trillion-parameter coding agents at scale isn't easy though: typically, each step launches thousands of rollouts, each with
Introducing SWE-2, our closest model yet to the frontier.
On leading evals, it scores on par with recent frontier models – at up to 70% lower cost.
We scaled RL to multiple trillions of parameters, with a refined recipe that pushes the Pareto curve on both capabilities & cost.
Your sandbox, your call.
Bring your own sandbox or connect a sandbox provider.
Choose the environment that fits your workload:
• CPU, GPU, and memory options
• Fully managed environments or deployments within your VPC
• File and secret storage options
We offer first-class
DeepSeek v4.1-Flash, a new multimodal model from @deepseek_ai, is now available on Modal.
v4.1-Flash uses DeepSeek's Causal Encoder-Decoder architecture, activating 16B parameters for input and 8B parameters for output to improve cost efficiency for input-heavy agentic