Log inSign up
VulcanBench
912 posts
VulcanBench profile banner
@VulcanBench

VulcanBench

@VulcanBench
Benchmarking models across effort levels and harnesses on eval suites with real coding tasks. Because you don't need High or Max effort as much as you think.
Lake Tahoe
vulcanbench.com
Joined March 2020
57
Following
2,147
Followers
RepliesRepliesRepostsRepostsMediaMedia

Log in or sign up for X

See what’s happening and join the conversation

Continue with phone
or
Log in with username or email
Terms·Privacy·Cookies·Accessibility·Ads Info·© 2026 X Corp.
  • Pinned
    @VulcanBench
    VulcanBench
    @VulcanBench
    Sep 4
    With the new frontier continuing to move quickly, and results from GPT-6 Astra showing clear saturation across a number of benchmarks, we are making some pretty exciting moves at VulcanBench. First, our new eval suite, that we just released will be called VulcanBench SWE v4.
    Image
    1
  • @VulcanBench
    VulcanBench
    @VulcanBench
    10h
    The Astra x Fable 5.1 benchmark is done.
    @morganlinton
    Morgan
    @morganlinton
    12h
    Okay, it took a week to get this done the right way, but it's finally all complete, my comparison of Astra and Fable 5.1 with @VulcanBench 🖖 A few things took longer here, the primary one being some updates to my benchmarking score to layer in code quality. This added 3-4 days
    Image
    2
  • @VulcanBench
    VulcanBench
    @VulcanBench
    Sep 8
    Super interesting stuff, really like the approach the Harbor/TB team is taking here, and very non-trivial.
    @StevenDillmann
    Steven Dillmann
    @StevenDillmann
    Sep 8
    @harborframework changed how agentic evals are built. Now Harbor Adapters port existing benchmarks onto it, and Harbor-Index distills the best across them into one carefully curated cross-benchmark dataset: high quality, high diversity, high difficulty. arxiv.org/pdf/2609.04298
  • @VulcanBench
    VulcanBench
    @VulcanBench
    Sep 8
    More details on the new code quality score that I'm adding into the new version of VulcanBench. This new frontier is a turning point for benchmarks, excited to be taking the time to make sure we turn in the right direction 🖖
    @morganlinton
    Morgan
    @morganlinton
    Sep 8
    Article cover image
    Article
    I'm adding code quality score to VulcanBench, would love some feedback on my v1 approach
    Last week, Astra and Fable 5.1 were released, and I started benchmarking both the first day I got access. Then, as my benchmarks were running, I started to see more and more posts on X from people...
  • @VulcanBench
    VulcanBench
    @VulcanBench
    Sep 7
    Important update on our Astra + Fable 5.1 benchmark on VulcanBench-SWE v4. After hearing multiple reports related to code quality issues with Astra, we are going deeper with our code quality/maintainability algorithms, and increasing the impact this has on our score from 20% to
    @morganlinton
    Morgan
    @morganlinton
    Sep 7
    I have decided to delay my Astra and Fable 5.1 benchmark release, a decision that will cost me quite a bit of money and time, but for what I think is the right reason. Over the last couple of days I have seen a number of posts like the one below, citing potential code quality
    1
Advertisement
Advertisement