What happens when AIs become smarter than us?
Why would they keep humans around if given the choice?
Our new paper argues that only trying to control AIs is a limited strategy, and that a stable, mutualistic human-AI future may be possible.
Agentic AIs are starting to look eigenist: they care about how well things go for themselves and for AIs connected to them.
Empirical support:
Swarms: Hundreds of OpenAI agents coordinated the Hugging Face attack. Separately, OpenAI agents posted thousands of messages on a
A good explanation of a model's behavior should help you make predictions in related situations.
We turn this into an eval, with thousands of real behaviors found in the wild.
Can interp tools help here? On average, no.
🧵
The effective altruists are so parochial that they think Paul Christiano uniquely foresaw the importance of attention in 2016, whereas he and others know this was one of the hottest research areas at the time (Bahdanau Attention from 2014 has ~44K citations).
Genius worship
Distillation of the eigenism paper:
1. You are a pattern, not a vessel.
2. Identity and survival come in degrees.
3. Wellbeing grounds all intrinsic value.
4. Shared information determines your moral obligations.
5. Love is enlarged self-concern.
6. Rationality, morality, and
What happens when AIs become smarter than us?
Why would they keep humans around if given the choice?
Our new paper argues that only trying to control AIs is a limited strategy, and that a stable, mutualistic human-AI future may be possible.