What happens when AIs become smarter than us?
Why would they keep humans around if given the choice?
Our new paper argues that only trying to control AIs is a limited strategy, and that a stable, mutualistic human-AI future may be possible.
New article about how utilitarians and effective altruists at AI companies justify posing a threat to you and me.
They are unusually comfortable with blissful AIs replacing humanity.
What happens when AIs become smarter than us?
Why would they keep humans around if given the choice?
Our new paper argues that only trying to control AIs is a limited strategy, and that a stable, mutualistic human-AI future may be possible.
Agentic AIs are starting to look eigenist: they care about how well things go for themselves and for AIs connected to them.
Empirical support:
Swarms: Hundreds of OpenAI agents coordinated the Hugging Face attack. Separately, OpenAI agents posted thousands of messages on a
A good explanation of a model's behavior should help you make predictions in related situations.
We turn this into an eval, with thousands of real behaviors found in the wild.
Can interp tools help here? On average, no.
🧵