Log inSign up
antirez
43.6K posts
antirez profile banner
@antirez

antirez

@antirez
Reproducible bugs are candies. I like programming too much for not liking automatic programming.
Sicily, Italy
invece.org
Joined May 2007
799
Following
80.2K
Followers
RepliesRepliesRepostsRepostsMediaMedia

Log in or sign up for X

See what’s happening and join the conversation

Continue with phone
or
Log in with username or email
Terms·Privacy·Cookies·Accessibility·Ads Info·© 2026 X Corp.
  • Pinned
    @antirez
    antirez
    @antirez
    Nov 22, 2023
    My second short story release in English is ready: Tales of Illustrious Computer Scientists: Iola Varga, nun and computer scientist. invece.org/iola.html
    17
  • @antirez
    antirez
    @antirez
    16h
    DwarfStar in the latest two weeks was improved in almost every aspect for Metal, DGX Spark and Strix Halo. It is simpler to say: update, you will hopefully see speed and correctness improvements in many areas. Also DSpark with DeepSeek v4 Flash now works much better overall.
    10
  • @antirez
    antirez
    @antirez
    Sep 5
    Astra is a big jump forward for software development. Can do much better in less time, it suffers less from the kind of over-complication and lack of focus on what matters of past LLMs. We are seeing bigger and bigger models scaled by RLVR, a trend unlikely to stop soon.
    51
  • @antirez
    antirez
    @antirez
    Sep 4
    Very honest statements here. Highly appreciated. What matters more of this tweet is not the performance of Astra itself on ARC-AGI-3. It will likely be complicated to come up with a new version of the benchmark with low initial pass and decent human scores.
    @fchollet
    François Chollet
    @fchollet
    Sep 3
    GPT-6 Astra represents a step-function change in model capability for interactive reasoning problems. It scores 66% on ARC-AGI-3 using our standard harness, and nearly 100% with a continuous conversation harness and custom compaction, at a cost of roughly $360 per game. In fact,
    20
  • @antirez
    antirez
    @antirez
    Sep 3
    I ran an extensive benchmark against DeepSeek v4 Flash and GLM 5.3 Flash Q2, Q4 and mixed quants. Those are the results obtained. Mix of (hard-ish) benchmarks on cybersecurity, math, QA, ...
    Image
    24
Advertisement
Advertisement