The easy coding problems are solved. The hard ones are the whole job now.
We're hiring senior engineers to find where frontier models break and set the bar they're trained against.
Remote, $100-150+/hr, no ML background needed.
Our mission is to raise AGI with the richness of humanity — curious, witty, imaginative, and full of breathtaking brilliance.
Joined June 2020
- We post-trained a model on office work. It also improved at coding. New paper from our research team! The training run: Office work RL environments (spreadsheets, documents, web research, planning). Zero coding tasks. The result: +5.8pp on SWE-Bench Pro. What transferred: A
- Most engineers will use AI. A few will get to shape it. We're hiring senior engineers to build the coding benchmarks frontier labs train on. Find where the best models break, set the bar they're measured against. Remote, contract, $100-150+/hr, no ML background needed.
- Frontier Data, Off the Shelf. The training data, evals, and RL environments we build for frontier labs, now available directly. ⌨️ Agentic coding: 35.2 → 47.2 on Terminal-Bench 2.0 🔧 Enterprise agents: 24.2 → 33.8 on Toolathlon 📐 STEM reasoning: 29.1 → 38.4 on
- Our Chartography benchmark measures whether models can read professional charts. @AnthropicAI’s contribution to science: what if it squinted? Nice job Fable.Replying to @ClaudeDevsOn the Chartography benchmark (100 questions over dense real-world charts), Fable 5's accuracy goes from 29% to 73% with a zoom tool, and Sonnet 5’s from 13% to 44%. See the cookbook to learn more: github.com/anthropics/cla…



