I wouldn't be surprised 1 year from now to have Opus 4.5 capabilities in a model a mortal human can run at home. By then sure, Opus 6.5++ will be available and way better, but it doesn't decrease the value proposition and price control effects of open source models. @Zai_org
Finally, I was able to reproduce Qwen's results on DeepSWE 1.1 with Claude Code for Qwen3.8 27B
Qwen published 42.2
I got 42.5
And it's a 7-point improvement over my runs that didn't preserve thinking!
Overall, my runs with Pi remain the best, reaching 46.
When I see the n-gram embedding lookup table in @Alibaba_Qwen Qwen-3.8-next I can't help but think this is a big step towards @karpathy 's cognitive core. I know they did some tests that showed 2 tables didn't help much, but I'm imagining ways to scale up these tables into an
For (1), agents modified their target programs to be easier to exploit & put the modified targets in cache. They then worked on crashing their targets in the hope that a restart would load the modified version from cache. Some agents risked failing their task to try this.