New creative writing leaderboard additions:
Qwen-3.8-2.4T
Muse-Glimmer-30B
NVIDIA-Nemotron-3-Ultra-550B
Gemini-3.6-flash
Inkling-Small
DeepSeek-V4-Flash-0731
Giving props to the overperformers: @Alibaba_Qwen for improving their previous rank by miles, and @Meta superintelligence
Announcing EQ-Bench 4
It's a 16-turn chat with a diverse cast of simulated user personas.
Core idea: There's no single best policy for user interaction.
Intuiting the user's preferences & needs is hard; getting it wrong means they lose trust or leave. We encode this challenge
Creative writing bench updates. We got:
- kimi-k3
- gpt-5.6
- muse-spark-1.1
- @thinkymachines first model Inkling
So much healthy competition, love to see it!
This is one of the last things I worked on at @liquidai, <3 that they open sourced it.
Antidoom generates preference data to punish the start of repetitive loops. It only adjusts the culprit tokens as far as needed, without damaging the model. It's effective! 23%->1% loop rate.
Today we release Antidoom, an open-source method that removes a common failure mode in reasoning models: the doom loop.
Doom-loop rates before and after, with eval scores up across the board:
> Early LFM2.5-2.6B checkpoint: 10.2% → 1.4%
> Qwen3.5-4B: 22.9% → 1% (greedy