My second short story release in English is ready: Tales of Illustrious Computer Scientists: Iola Varga, nun and computer scientist. invece.org/iola.html
DwarfStar in the latest two weeks was improved in almost every aspect for Metal, DGX Spark and Strix Halo. It is simpler to say: update, you will hopefully see speed and correctness improvements in many areas. Also DSpark with DeepSeek v4 Flash now works much better overall.
Astra is a big jump forward for software development. Can do much better in less time, it suffers less from the kind of over-complication and lack of focus on what matters of past LLMs. We are seeing bigger and bigger models scaled by RLVR, a trend unlikely to stop soon.
Very honest statements here. Highly appreciated. What matters more of this tweet is not the performance of Astra itself on ARC-AGI-3. It will likely be complicated to come up with a new version of the benchmark with low initial pass and decent human scores.
GPT-6 Astra represents a step-function change in model capability for interactive reasoning problems. It scores 66% on ARC-AGI-3 using our standard harness, and nearly 100% with a continuous conversation harness and custom compaction, at a cost of roughly $360 per game.
In fact,
I ran an extensive benchmark against DeepSeek v4 Flash and GLM 5.3 Flash Q2, Q4 and mixed quants. Those are the results obtained. Mix of (hard-ish) benchmarks on cybersecurity, math, QA, ...