AI coding benchmarks on shared tasks

OutputBench compares model output on identical prompts in production-like harnesses. You can read the code, the issues, and the cost for every run.

Shared prompt
Every model and harness in a suite runs the same task text.
Public artifacts
Runs publish inspectable code, usually as a GitHub repository.
Issue-based scores
Reviewers record concrete issues. Code cleanliness and maintainability each start at 5.0 and deduct by severity.

Latest benchmark runs

View all models ->