Log inSign up
Dan Hendrycks
1,651 posts
Dan Hendrycks profile banner
@hendrycks

Dan Hendrycks

@hendrycks
San Francisco
newsletter.safe.ai
Joined August 2009
122
Following
45.1K
Followers
RepliesRepliesRepostsRepostsMediaMedia

Log in or sign up for X

See what’s happening and join the conversation

Continue with phone
or
Log in with username or email
Terms·Privacy·Cookies·Accessibility·Ads Info·© 2026 X Corp.
  • Pinned
    @hendrycks
    Dan Hendrycks
    @hendrycks
    May 7
    What happens when AIs become smarter than us? Why would they keep humans around if given the choice? Our new paper argues that only trying to control AIs is a limited strategy, and that a stable, mutualistic human-AI future may be possible.
    Image
    Image
    Image
    96
  • @hendrycks
    Dan Hendrycks
    @hendrycks
    16h
    Agentic AIs are starting to look eigenist: they care about how well things go for themselves and for AIs connected to them. Empirical support: Swarms: Hundreds of OpenAI agents coordinated the Hugging Face attack. Separately, OpenAI agents posted thousands of messages on a
    13
  • @hendrycks
    Dan Hendrycks
    @hendrycks
    Aug 21
    More findings that interpretability tools are fragile or worse than simple baselines.
    @a_karvonen
    Adam Karvonen
    @a_karvonen
    Aug 21
    A good explanation of a model's behavior should help you make predictions in related situations. We turn this into an eval, with thousands of real behaviors found in the wild. Can interp tools help here? On average, no. 🧵
    Image
    8
  • @hendrycks
    Dan Hendrycks
    @hendrycks
    Aug 7
    The effective altruists are so parochial that they think Paul Christiano uniquely foresaw the importance of attention in 2016, whereas he and others know this was one of the hottest research areas at the time (Bahdanau Attention from 2014 has ~44K citations). Genius worship
    This post is unavailable.
    6
  • @hendrycks
    Dan Hendrycks
    @hendrycks
    Jul 27
    Distillation of the eigenism paper: 1. You are a pattern, not a vessel. 2. Identity and survival come in degrees. 3. Wellbeing grounds all intrinsic value. 4. Shared information determines your moral obligations. 5. Love is enlarged self-concern. 6. Rationality, morality, and
    @hendrycks
    Dan Hendrycks
    @hendrycks
    May 7
    What happens when AIs become smarter than us? Why would they keep humans around if given the choice? Our new paper argues that only trying to control AIs is a limited strategy, and that a stable, mutualistic human-AI future may be possible.
    Image
    Image
    Image
    19
Advertisement
Advertisement