Announcing EQ-Bench 4
It's a 16-turn chat with a diverse cast of simulated user personas.
Core idea: There's no single best policy for user interaction.
Intuiting the user's preferences & needs is hard; getting it wrong means they lose trust or leave. We encode this challenge
- Creative writing bench updates. We got: - kimi-k3 - gpt-5.6 - muse-spark-1.1 - @thinkymachines first model Inkling So much healthy competition, love to see it!
- This is one of the last things I worked on at @liquidai, <3 that they open sourced it. Antidoom generates preference data to punish the start of repetitive loops. It only adjusts the culprit tokens as far as needed, without damaging the model. It's effective! 23%->1% loop rate.
- Fable 5 tops EQ-Bench and both creative writing evals! Personal take from reading some outputs: It has tics and tells, and isn't compelling in the way human writing is. But I think it earned its spot, in the sense that it's *relatively* excellent & hasn't reward-hacked the eval.


