Official Leaderboards
Each entry reports % Resolved, the percentage of task instances solved.
- Open-weights model
- Run performed or directly checked by the SWE-bench team
SWE-bench
2294 instancesThe original benchmark: real GitHub issues from 12 Python repositories.
Read the postBash Only
500 instancesThe default Verified view: every model in the same mini-SWE-agent environment.
Read the postNews
-
May 2026
We released ProgramBench to benchmark whether models can code meaningful software artifacts from scratch.
Link
-
Nov 2025 Introducing CodeClash, our new eval of LMs as goal (not task) oriented developers. Link
-
Jul 2025 mini-SWE-agent scores 65% on SWE-bench Verified in 100 lines of Python. Link
-
May 2025 SWE-smith is out! Train your own models for software engineering agents. Link
-
Mar 2025 SWE-agent 1.0 is the open source SOTA on SWE-bench Lite. Link
-
Oct 2024 Introducing SWE-bench Multimodal. Link
-
Aug 2024 SWE-bench x OpenAI = SWE-bench Verified. Report
-
Jun 2024 Docker-ized SWE-bench for easier evaluation. Report
-
Mar 2024 Check out SWE-agent (12.47% on SWE-bench). Link
-
Mar 2024 Released SWE-bench Lite. Report
Acknowledgements
We thank the following institutions for their generous support: Open Philanthropy, AWS, Modal, Andreessen Horowitz, OpenAI, and Anthropic.