Skip to content

[codex] docs(eval): publish nine-arm Terminal-Bench 2.1 leaderboard - #3004

Draft
hqhq1025 wants to merge 1 commit into
apache:mainfrom
hqhq1025:codex/tb21-nine-arm-report
Draft

[codex] docs(eval): publish nine-arm Terminal-Bench 2.1 leaderboard#3004
hqhq1025 wants to merge 1 commit into
apache:mainfrom
hqhq1025:codex/tb21-nine-arm-report

Conversation

@hqhq1025

Copy link
Copy Markdown
Contributor

Summary

Docs-only: publish the final nine-harness Terminal-Bench 2.1 comparison around DeepSeek V4 Flash at max reasoning effort, plus an 89-task accepted-outcome CSV.

This extends the four-arm report in #2208 with OpenCode, Kimi Code, ZCode, Pi, and DeepSeek Harness (DSH).

Full report: docs/eval/terminal-bench-2.1-deepseek-v4-flash-nine-arm.md

Leaderboard

Rank Harness Passed Pass@1 Tokens Avg/task Cache Cost
1 Codex 73/89 82.0% 259.62M 2.92M 98.6% $2.105
2 Maka 69/89 77.5% 185.71M 2.09M 98.8% $1.815
3 Pi 66/89 74.2% 131.40M 1.48M 98.5% $1.303
4 DSH 65/89 73.0% 182.84M 2.05M 98.6% $1.705
5 ZCode 63/89 70.8% 223.46M 2.51M 98.9% $2.287
6 Reasonix 60/89 67.4% 216.17M 2.43M 99.0% $1.935
7 OpenCode 58/89 65.2% 166.44M 1.87M 98.6% $1.745
8 Kimi Code 53/89 59.6% 193.75M 2.18M 98.2% $2.162
9 Claude Code 49/89 55.1% 247.60M 2.78M 98.9% $2.633

All nine harnesses pass the same 28/89 tasks. Five tasks fail on all nine.

Scope and accounting

  • Terminal-Bench 2.1 revision d49e28f1e4ddd13d289e85a5f312a66750951932
  • 89 tasks, one repetition, official verifier, task-native timeout ×1
  • deepseek-v4-flash, thinking enabled, reasoning effort max
  • web tools removed from the provider-visible surface; shell networking retained under benchmark-contamination egress filtering
  • the original eight arms ran as one cohort; DSH is a configuration-aligned later run and is disclosed as descriptive rather than simultaneously paired
  • recorded token, cache, and cost aggregates are reported directly without extrapolation

DSH recovery disclosure

DSH initially recorded 61/89. Real-machine diagnosis found three Eval-specific defects: a five-minute Bash timeout that interrupted dpkg, interactive tzdata setup, and background-service teardown before verifier execution. Commit 28ebe0949 fixed those paths.

Seven cells were rerun. Four changed from fail to pass (hf-model-inference, merge-diff-arc-agi-task, polyglot-c-py, regex-log), producing the final 65/89 result. Three remained genuine verifier failures.

Verification

  • CSV contains exactly 89 rows and 89 unique task ids
  • recomputed pass totals: 73, 69, 66, 65, 63, 60, 58, 53, 49
  • recomputed agreement: 28 all-pass, 5 all-fail, 56 mixed
  • CSV SHA-256: c27c3bcbfc3ebe8e21cc250dc409f02f49ae055032eaf2f19fd6349986e94e6a
  • git diff --check passed
  • secret-shaped token scan passed
  • docs-only change; repository code tests were not run

Review focus

  • whether the independent DSH run is disclosed strongly enough relative to the simultaneous eight-arm cohort
  • whether the accepted-attempt and recovery policy is clear
  • whether the recorded-economics convention is described without implying billing precision

@M4n5ter
M4n5ter force-pushed the codex/tb21-nine-arm-report branch from 40045ff to 63a861c Compare August 26, 2026 10:03
@github-actions github-actions Bot added the effort/M Under 500 readable lines label Aug 27, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

effort/M Under 500 readable lines

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant