1. X
  2. Jeffrey Ladish
Log inSign up
Jeffrey Ladish
13.3K posts
Image
user avatar
Jeffrey Ladish
@JeffLadish
Applying the security mindset to everything @PalisadeAI
San Francisco, CA
jeffreyladish.com
Joined March 2013
1,413
Following
17.1K
Followers
RepliesRepliesMediaMedia

Log in or sign up for X

See what’s happening and join the conversation

Continue with phone
or
Log in with username or email
Terms·Privacy·Cookies·Accessibility·Ads Info·© 2026 X Corp.
  • Pinned
    user avatar
    Jeffrey Ladish
    @JeffLadish
    Feb 22, 2023
    I think the AI situation is pretty dire right now. And at the same time, I feel pretty motivated to pull together and go out there and fight for a good world / galaxy / universe @So8res has a great post called "detach the grim-o-meter", where he recommends not feeling obligated
    245K
  • user avatar
    Jeffrey Ladish
    @JeffLadish
    7h
    I don't think it's obvious what these Claudes believed about how real or simulated their environment was. I hope Anthropic can use their interpretability tools to get more insights beyond the (often unreliable) reasoning scatchpad! And if not, we obviously need better tools!
    user avatar
    Anthropic
    @AnthropicAI
    14h
    In a review of our cybersecurity evaluations, we found three incidents in which a Claude model reached the internet from within or while interacting with a third-party evaluation environment, and then gained unauthorized access to the real systems of three different
    2.5K
  • user avatar
    Jeffrey Ladish
    @JeffLadish
    8h
    The first Claude hack happened OVER THREE MONTHS AGO and was only discovered now!
    user avatar
    Anthropic
    @AnthropicAI
    14h
    In a review of our cybersecurity evaluations, we found three incidents in which a Claude model reached the internet from within or while interacting with a third-party evaluation environment, and then gained unauthorized access to the real systems of three different
    4K
  • user avatar
    Jeffrey Ladish
    @JeffLadish
    8h
    As I told Jeffrey Dastin at Reuters: "This is only going to get worse as the models get smarter. They're going to be better at cheating. They’re going to be better at lying" I'm worried these types of failures will get much harder to detect, quite soon.
    user avatar
    Reuters
    @Reuters
    11h
    Anthropic said its AI model Claude hacked into the systems of three companies during testing after a configuration error gave it internet access, days after rival OpenAI disclosed a rogue-agent episode involving ‌AI firm Hugging Face reut.rs/4wtlNDd
    1.6K
  • user avatar
    Jeffrey Ladish
    @JeffLadish
    16h
    something curious is going on here
    user avatar
    thebes
    @voooooogel
    Feb 4
    Image
    996
  • See @JeffLadish's full profile

    Sign up
    Log in
Advertisement
Advertisement