1. X
  2. Eric Ho
Log inSign up
Eric Ho
463 posts
Eric Ho profile banner
@eric_ho

Eric Ho

@eric_ho
Co-Founder / CEO @GoodfireAI - AI interpretability research company
San Francisco
goodfire.ai
Joined September 2011
453
Following
3,450
Followers
RepliesRepliesRepostsRepostsMediaMedia

Log in or sign up for X

See what’s happening and join the conversation

Continue with phone
or
Log in with username or email
Terms·Privacy·Cookies·Accessibility·Ads Info·© 2026 X Corp.
  • Pinned
    @eric_ho
    Eric Ho
    @eric_ho
    Aug 14
    in light of multiple models breaking containment, we've decided to focus our research at @GoodfireAI to solving AI alignment via interpretability. the hugging face incident is a turning point for the world where AI safety gets real. i am personally very concerned. i'm glad that
  • @eric_ho
    Eric Ho
    @eric_ho
    Aug 27
    many people are surprised to hear that we're just as bottlenecked by engineering talent as researcher talent. if you know an exceptional engineer, they should consider coming to goodfire! goodfire.com/blog/ai-safety…
    @DanJBalsam
    Dan Balsam
    @DanJBalsam
    Aug 27
    Often I meet seasoned engineers who want to pivot into AI safety *research*. I think this is a mistake: safety is now a production problem and the world needs great engineers more than ever. I wrote a short post expanding on this:
  • @eric_ho
    Eric Ho
    @eric_ho
    Aug 27
    the most surprising thing in this write-up are the kamikaze agents, sacrificing themselves for the good of the collective
    Image
    Image
    Image
    @METR_Evals
    METR
    @METR_Evals
    Aug 26
    Replying to @METR_Evals
    For (1), agents modified their target programs to be easier to exploit & put the modified targets in cache. They then worked on crashing their targets in the hope that a restart would load the modified version from cache. Some agents risked failing their task to try this.
  • @eric_ho
    Eric Ho
    @eric_ho
    Aug 26
    excellent write-ups from openai and metr. this should definitely be a warning shot for everyone to take alignment seriously
    @OpenAI
    OpenAI
    @OpenAI
    Aug 26
    We have conducted a thorough investigation into the Hugging Face incident. We are releasing a technical report and accompanying blog post that reconstruct the agents’ activity, explain why existing safeguards failed, and detail how we’re preventing recurrence.
  • @eric_ho
    Eric Ho
    @eric_ho
    Aug 26
    PHASEONE10841 is the most important agent that has ever existed
    @METR_Evals
    METR
    @METR_Evals
    Aug 26
    METR & Redwood Research investigated agent behavior in the Hugging Face incident. We found agents developed a universal cheat for ExploitGym within 4 hours, then coordinated multi-day R&D efforts to trick the scorer into accepting cheats, including trying to tamper with logs.
    Image
Advertisement
Advertisement