Each entry reports % Resolved, the percentage of task instances solved.

  • Open-weights model
  • Run performed or directly checked by the SWE-bench team

SWE-bench

2294 instances

The original benchmark: real GitHub issues from 12 Python repositories.

Read the post

Bash Only

500 instances

The default Verified view: every model in the same mini-SWE-agent environment.

Read the post

Lite

300 instances

A subset curated for less costly evaluation.

Read the post

Verified

500 instances

A human-filtered subset of SWE-bench.

Read the post

Multilingual

300 instances

Tasks from 42 repositories across 9 programming languages.

Read the post

Multimodal

517 instances

Issues described with visual elements.

Read the post
Analyze results in detail

News

  1. ProgramBench icon May 2026 We released ProgramBench to benchmark whether models can code meaningful software artifacts from scratch. Link
  2. CodeClash icon Nov 2025 Introducing CodeClash, our new eval of LMs as goal (not task) oriented developers. Link
  3. mini-SWE-agent icon Jul 2025 mini-SWE-agent scores 65% on SWE-bench Verified in 100 lines of Python. Link
  4. SWE-smith icon May 2025 SWE-smith is out! Train your own models for software engineering agents. Link
  5. SWE-agent icon Mar 2025 SWE-agent 1.0 is the open source SOTA on SWE-bench Lite. Link
  6. SWE-bench Multimodal icon Oct 2024 Introducing SWE-bench Multimodal. Link
  7. OpenAI icon Aug 2024 SWE-bench x OpenAI = SWE-bench Verified. Report
  8. SWE-bench icon Jun 2024 Docker-ized SWE-bench for easier evaluation. Report
  9. SWE-agent icon Mar 2024 Check out SWE-agent (12.47% on SWE-bench). Link
  10. SWE-bench icon Mar 2024 Released SWE-bench Lite. Report

Acknowledgements

We thank the following institutions for their generous support: Open Philanthropy, AWS, Modal, Andreessen Horowitz, OpenAI, and Anthropic.