We're open-sourcing Claude Commerce Agents.
This is a blueprint for building shopping and merchant agents, with reference implementations across retail, travel, telecom, and entertainment.
Fable gets a glow up. SOTA on the benchmarks, very excellent on low reasoning, and with a 75% drop on api cache reads for those long running agentic workloads
Across our benchmarks, the model sets a new standard.
It scores 52.6% on Terminal-Bench-Science 0.1, more than double Fable 5. On Terminal-Bench 4.0, it scores 55.8% against 42.0% for Fable 5.
“Harness” does not mean application. An application can have many harnesses. A harness is a loop that runs the model. Its purpose is to get the model to give you a good answer. The application uses the harness(es) to give you a good experience.