Pinned
Check out the Pareto frontier of success rate vs cost on terminal tasks and more.
p.s. ox-alpha (GLM-5.3 Flash) is completely dominated by GPT-5.6 Luna and DeepSeek V4 Flash 🙃
Newly released models repeatedly appear near the top on the @terminalbench 2.1 leaderboard. Are these models actually on par with frontier models on terminal tasks?
We took the top 20 models from that leaderboard and ran them on TB-fn. TB-fn reworks the same 89 tasks by adding





