Log inSign up
Dan Hendrycks
1,656 posts
Dan Hendrycks profile banner
@hendrycks

Dan Hendrycks

@hendrycks
eigenism.org
San Francisco
newsletter.safe.ai
Joined August 2009
122
Following
45.3K
Followers
RepliesRepliesRepostsRepostsMediaMedia

Log in or sign up for X

See what’s happening and join the conversation

Continue with phone
or
Log in with username or email
Terms·Privacy·Cookies·Accessibility·Ads Info·© 2026 X Corp.
  • Pinned
    @hendrycks
    Dan Hendrycks
    @hendrycks
    May 7
    What happens when AIs become smarter than us? Why would they keep humans around if given the choice? Our new paper argues that only trying to control AIs is a limited strategy, and that a stable, mutualistic human-AI future may be possible.
    Image
    Image
    Image
    98
  • @hendrycks
    Dan Hendrycks
    @hendrycks
    3h
    New article about how utilitarians and effective altruists at AI companies justify posing a threat to you and me. They are unusually comfortable with blissful AIs replacing humanity.
    Image
    Suicidal Compassion: How Utilitarianism at AI Companies Endangers Humanity | AI Frontiers
    From ai-frontiers.org
    3
  • @hendrycks
    Dan Hendrycks
    @hendrycks
    Sep 8
    “AI makes philosophy honest.” - Daniel Dennett
    @hendrycks
    Dan Hendrycks
    @hendrycks
    May 7
    What happens when AIs become smarter than us? Why would they keep humans around if given the choice? Our new paper argues that only trying to control AIs is a limited strategy, and that a stable, mutualistic human-AI future may be possible.
    Image
    Image
    Image
    5
  • @hendrycks
    Dan Hendrycks
    @hendrycks
    Sep 6
    Agentic AIs are starting to look eigenist: they care about how well things go for themselves and for AIs connected to them. Empirical support: Swarms: Hundreds of OpenAI agents coordinated the Hugging Face attack. Separately, OpenAI agents posted thousands of messages on a
    16
  • @hendrycks
    Dan Hendrycks
    @hendrycks
    Aug 21
    More findings that interpretability tools are fragile or worse than simple baselines.
    @a_karvonen
    Adam Karvonen
    @a_karvonen
    Aug 21
    A good explanation of a model's behavior should help you make predictions in related situations. We turn this into an eval, with thousands of real behaviors found in the wild. Can interp tools help here? On average, no. 🧵
    Image
    8
Advertisement
Advertisement