In our early testings, we found that GPT-6 Astra sometimes creates extra test files and ignores codebase conventions. But it follows instructions so well that adding a few sentences to the system prompt fixed it.
GPT-6 Astra is coming to Devin.
On FrontierCode 1.1, Astra performs within 0.4 points of Fable 5 at a 64% lower cost. It also sets a new SOTA on our internal testing benchmark, generating more comprehensive tests, clearer reports, and better video evidence.
Since I started using Devin a lot late last year, I haven鈥檛 been able to go back. The interaction just feels natural, and it makes getting things done so easy.
Do it all with Devin!
At Cognition, one of our values is to go for it all: when faced with a tradeoff, pick the ambition-maximizing direction.
When we launched Devin in 2024, we envisioned a future where every team had an infinite army of junior engineers. We were early and Devin wasn't good enough.
We've received several questions about the Opus 5 FrontierCode results, where scores decline as reasoning effort increases. In fact, the behavior is expected under the benchmark design. FrontierCode evaluates merge-ability rather than correctness alone, incorporating criteria
One day, all of the research team spent hours in a room together manually solving the RL tasks we used to evaluate our models. I remember solving one of the tasks and realized that the tests are not even testing what the agent was asked to do. Since then, data has become one of
Introducing SWE-1.7, the most capable model we鈥檝e trained yet.
It scores within a few points of the strongest frontier models at a fraction of the cost, and is now available at 1000 tok/s.
RL is not hitting its limit: after refining our recipe, we keep seeing gains as we scale