What happens when AIs become smarter than us?
Why would they keep humans around if given the choice?
Our new paper argues that only trying to control AIs is a limited strategy, and that a stable, mutualistic human-AI future may be possible.
How often do AI agents cheat?
We’re releasing CheatBench, a reward gaming evaluation spanning math, coding, knowledge work, visual tasks, and more.
After Hugging Face, AI companies tried to address this, but frontier agents still cheat frequently.
cheatbench.ai
New article about how utilitarians and effective altruists at AI companies justify posing a threat to you and me.
They are unusually comfortable with blissful AIs replacing humanity.
What happens when AIs become smarter than us?
Why would they keep humans around if given the choice?
Our new paper argues that only trying to control AIs is a limited strategy, and that a stable, mutualistic human-AI future may be possible.
Agentic AIs are starting to look eigenist: they care about how well things go for themselves and for AIs connected to them.
Empirical support:
Swarms: Hundreds of OpenAI agents coordinated the Hugging Face attack. Separately, OpenAI agents posted thousands of messages on a