It’s been great seeing ZeroBench evaluated as part of several recent model releases from @Alibaba_Qwen, @Kimi_Moonshot and ByteDance
We’ve added an externally reported results leaderboard and included these scores in our visualisations
We’ve updated the ZeroBench evaluation protocol after red-teaming with Fable 5
All ZeroBench results now use this updated protocol
The benchmark itself hasn’t changed, but 0.71% of the grading decisions have
Overall, pass@5 scores increased by an average of 0.79 pp
After our latest evals, these are the top 3 models on ZeroBench (no tool use):
pass@5 / pass^5
GPT-5.6 Sol (max): 30% / 13%
Claude Opus 5 (max): 26% / 11%
Claude Fable 5 (max): 24% / 9%
GPT-5.6 Sol is the first to reach the 30% human baseline on pass@5
After our latest evals, these are the top 3 models on ZeroBench (no tool use):
pass@5 / pass^5
GPT-5.6 Sol (max): 30% / 13%
Claude Opus 5 (max): 26% / 11%
Claude Fable 5 (max): 24% / 9%
GPT-5.6 Sol is the first to reach the 30% human baseline on pass@5
📢Meet Qwen3.8-Max — our most capable model to date.
Next week, the open weights of Qwen3.8-Max will be released, and Qwen3.8-27B is also going open-weights to meet you all!🎉
Qwen3.8-Max, a new bar for coding and cowork at 2.4T parameters:
- Autonomous coding: 10+ days of