Measuring AI models on Git tasks

Benchmarks for git-related tasks across a variety of LLMs. Pick between coding assistants for the right tasks. Learn more about the motivations for GitBench on the blog. Read the Blog ->

Get the GitBench analysis PDF

Enter your email and we'll send you the PDF with our notes on what the results mean.

The aggregated and averaged results for each model across all benchmarks and effort levels. The whiskers show the range of values from the effort levels to give you an idea of the range of results.

TextJSON

Pick two criteria. The shaded quadrant is the optimal direction. For example, lower cost with higher pass rate.

Loading...

Input and Output tokens as well as Reasoning tokens for models that support it.

Loading...

API call latency across successful fixture calls. It excludes fixture setup, scoring, and cleanup.

Loading...

Source token counts come from provider-reported usage; GitBench calculates the displayed aggregates from them. Provider tokenizers and accounting methods may differ. Learn more ->

Loading...

Rows are Git task categories, columns are models. Green cells show where a model excels. Red cells reveal weaknesses.

Loading...