1. X
  2. Sam Paech
Log inSign up
Sam Paech
1,208 posts
Image
user avatar
Sam Paech
@sam_paech
Building @thinkymachines Maintainer of EQ-Bench github.com/orgs/EQ-bench/… github.com/sam-paech
San Francisco
eqbench.com
Joined July 2012
237
Following
4,137
Followers
RepliesRepliesMediaMedia
  • user avatar
    Sam Paech
    @sam_paech
    Jul 23
    Announcing EQ-Bench 4 It's a 16-turn chat with a diverse cast of simulated user personas. Core idea: There's no single best policy for user interaction. Intuiting the user's preferences & needs is hard; getting it wrong means they lose trust or leave. We encode this challenge
    Image
    Image
    Image
    Image
    14K
  • user avatar
    Sam Paech
    @sam_paech
    Jul 18
    Creative writing bench updates. We got: - kimi-k3 - gpt-5.6 - muse-spark-1.1 - @thinkymachines first model Inkling So much healthy competition, love to see it!
    Image
    Image
    62K
  • user avatar
    Sam Paech
    @sam_paech
    Jul 7
    This is one of the last things I worked on at @liquidai, <3 that they open sourced it. Antidoom generates preference data to punish the start of repetitive loops. It only adjusts the culprit tokens as far as needed, without damaging the model. It's effective! 23%->1% loop rate.
    user avatar
    Liquid AI
    @liquidai
    Jul 7
    Today we release Antidoom, an open-source method that removes a common failure mode in reasoning models: the doom loop. Doom-loop rates before and after, with eval scores up across the board: > Early LFM2.5-2.6B checkpoint: 10.2% → 1.4% > Qwen3.5-4B: 22.9% → 1% (greedy
    Image
    GIF
    5.7K
  • user avatar
    Sam Paech
    @sam_paech
    Jun 17
    GLM-5.2 is an extremely strong open weights model. Kudos to the @Zai_org team.
    Image
    Image
    Image
    Image
    7.3K
  • user avatar
    Sam Paech
    @sam_paech
    Jun 10
    Fable 5 tops EQ-Bench and both creative writing evals! Personal take from reading some outputs: It has tics and tells, and isn't compelling in the way human writing is. But I think it earned its spot, in the sense that it's *relatively* excellent & hasn't reward-hacked the eval.
    Image
    Image
    Image
    8.5K
  • See @sam_paech's full profile

    Sign up
    Log in

Log in or sign up for X

See what’s happening and join the conversation

Continue with phone
or
Log in with username or email
Terms·Privacy·Cookies·Accessibility·Ads Info·© 2026 X Corp.
Advertisement
Advertisement