The first ProgramBench task was just solved by GPT 5.5 high/xhigh. Interestingly, high/xhigh picked two different languages for the task (C vs Python). GPT 5.5 xhigh was significantly better than Opus 4.7 xhigh in all metrics. 🧵
In 2026, debugging GitHub issues communicated w/ visuals (not text) remains partially unsolved.
Since Opus 4.7, Anthropic models have steadily climbed in Multimodal performance.
Getting to 90+% with this benchmark is just a matter of time.
Had a great time talking at @ScaleAILabs. New levels of model capabilities require new benchmarking paradigms.
Also really enjoyed talks by @sirrice@mohit_r9a and @MiguelR33478246!
Thanks to all the organizers and to @liuying04
Extremely cool project: What's the best SWE-bench score you can achieve by training a model from scratch with limited $ budget. SWE-bench might be near saturated for frontier models, but it still has a lot of life in it for hillclimbing with smaller or from scratch models
We turned @karpathy's nanochat into a SWE-bench speedrun!
⚡ $60 of train compute → 5.0% pass@1 ≈ Claude 2 (2023 SOTA)
⚡ $1000 → 11.0% pass@1 > Claude 3 Haiku
From randomly initialized model weights!
How? 🧵👇