Log inSign up
Buck Shlegeris
1,114 posts
Buck Shlegeris profile banner
@bshlgrs

Buck Shlegeris

@bshlgrs
CEO@Redwood Research (@redwood_ai), working on technical research to reduce catastrophic risk from AI misalignment. [email protected]
Berkeley, CA
redwoodresearch.org
Joined January 2015
362
Following
6,068
Followers
RepliesRepliesRepostsRepostsMediaMedia

Log in or sign up for X

See what’s happening and join the conversation

Continue with phone
or
Log in with username or email
Terms·Privacy·Cookies·Accessibility·Ads Info·© 2026 X Corp.
  • Pinned
    @bshlgrs
    Buck Shlegeris
    @bshlgrs
    Apr 16, 2025
    We’ve just released the biggest and most intricate study of AI control to date, in a command line agent setting. IMO the techniques studied are the best available option for preventing misaligned early AGIs from causing sudden disasters, e.g. hacking servers they’re working on.
    9
  • @bshlgrs
    Buck Shlegeris
    @bshlgrs
    Sep 2
    I am extremely concerned by the reporting that Astra uses opaque recurrence. I don’t know whether Astra is much less CoT monitorable than previous models. But if OpenAI pushes this technique further, they’ll have the option to massively increase the recurrence and totally
    24
  • @bshlgrs
    Buck Shlegeris
    @bshlgrs
    Aug 26
    I’m very proud of Ryan, Ajeya, and Hjalmar’s work on this report. They did a great job of investigating this with very limited time. I hope that this strengthens the growing precedent of AI companies working with third party investigators to study misalignment incidents.
    @METR_Evals
    METR
    @METR_Evals
    Aug 26
    METR & Redwood Research investigated agent behavior in the Hugging Face incident. We found agents developed a universal cheat for ExploitGym within 4 hours, then coordinated multi-day R&D efforts to trick the scorer into accepting cheats, including trying to tamper with logs.
    Image
    5
  • @bshlgrs
    Buck Shlegeris
    @bshlgrs
    Aug 12
    I'm excited for Em and team's work on measuring AI conceptual reasoning capabilities!
    @emwcooper
    Emery Cooper
    @emwcooper
    Aug 12
    We want AIs to be able to help with work to reduce AI risk. But while models do great in domains where reliable feedback is relatively cheap and abundant, like Math and coding, a lot of work on AI risk isn't like that. Instead, we have to rely on good argumentation to answer
    Image
    2
  • @bshlgrs
    Buck Shlegeris
    @bshlgrs
    Aug 7
    I regret saying this. If AI developers competently implement safety measures we know about, risk from sub-ASI misalignment will be way lower. But these techniques probably fail for superintelligence. And it's very unclear whether better techniques will be developed in time.
    @GarrisonLovely
    Garrison Lovely
    @GarrisonLovely
    Aug 7
    Thinking about this quote from @redwood_ai director @bshlgrs, one of the pioneers of the field of AI control.
    Image
    21
Advertisement
Advertisement