1. X
  2. Ahmad Beirami
Log inSign up
Ahmad Beirami
4,964 posts
@abeirami

Ahmad Beirami

@abeirami
co-founder / ceo @fidian
sf | nyc
fidian.ai
Joined December 2018
3,158
Following
12.2K
Followers
1
Subscription
RepliesRepliesRepostsRepostsMediaMediaArticlesArticles

Log in or sign up for X

See what’s happening and join the conversation

Continue with phone
or
Log in with username or email
Terms·Privacy·Cookies·Accessibility·Ads Info·© 2026 X Corp.
  • Pinned
    @abeirami
    Ahmad Beirami
    @abeirami
    Aug 26
    Check out the Pareto frontier of success rate vs cost on terminal tasks and more. p.s. ox-alpha (GLM-5.3 Flash) is completely dominated by GPT-5.6 Luna and DeepSeek V4 Flash 🙃
    Image
    Image
    00:30
    @fidian
    Fidian
    @fidian
    Aug 26
    Newly released models repeatedly appear near the top on the @terminalbench 2.1 leaderboard. Are these models actually on par with frontier models on terminal tasks? We took the top 20 models from that leaderboard and ran them on TB-fn. TB-fn reworks the same 89 tasks by adding
  • @abeirami
    Ahmad Beirami
    @abeirami
    Aug 28
    👀
    @enginoid
    Fred Jonsson
    @enginoid
    Aug 28
    Replying to @abeirami
    Interesting -- we desperately need scalable ways to reliably analyze trajectories!
  • @abeirami
    Ahmad Beirami
    @abeirami
    Aug 27
    Suspicious! We had found that some models exhibit significant pass@3 drops when you replace the tasks with variants that have not been publicly released, and their ranking can move significantly. Other models remain stable. One would wonder why. 🤔 fidian.ai/blog/tb-fn-ben…
    Image
    Image
    Image
    @zainhas
    Zain
    Together AI
    @zainhas
    Aug 27
    there's no way this is passes vibe checks... but peep the multipass chart below for GLM-5.3 Flash and Fable 5 on DeepSWE
  • @abeirami
    Ahmad Beirami
    @abeirami
    Aug 27
    Reminder that $/token is an arbitrary and meaningless metric to internalize cost. Opus 5 (max) uses **five** times more tokens than GPT 5.6 Sol (max) to solve similar terminal tasks at a similar accuracy level! $/attempt on different types of tasks is much more meaningful.
  • @abeirami
    Ahmad Beirami
    @abeirami
    Aug 26
    The frontier on terminal agents looks very different once we (1) evaluate harder variants of the same task families and (2) fix the verifiers so they stop rewarding hacks and stop failing otherwise correct solutions for criteria that were never stated.
    @fidian
    Fidian
    @fidian
    Aug 26
    Newly released models repeatedly appear near the top on the @terminalbench 2.1 leaderboard. Are these models actually on par with frontier models on terminal tasks? We took the top 20 models from that leaderboard and ran them on TB-fn. TB-fn reworks the same 89 tasks by adding
    Image
    00:00
Advertisement
Advertisement