All three GPT-6 models sit on the ZeroBench performance-cost pareto frontier
pass@5 | pass^5
GPT-6 Astra (max)
51% | 35%
GPT-6 Sol (max)
41% | 17%
GPT-6 Luna (max)
20% | 7%
GPT-6 Astra (max) on ZeroBench:
pass@5: 51% (prev. SOTA: 30%)
pass^5: 35% (prev. SOTA: 13%)
If this reflects a genuine capability improvement, it is a significant step forward in both performance and consistency