1. X
  2. Center for AI Safety
Log inSign up
Center for AI Safety
319 posts
Image
user avatar
Center for AI Safety
@CAIS
Reducing societal-scale risks from AI.
San Francisco
safe.ai
Joined August 2022
3
Following
10.6K
Followers
RepliesRepliesArticlesArticlesMediaMedia
  • Pinned
    user avatar
    Center for AI Safety
    @CAIS
    May 30, 2023
    We’ve released a statement on the risk of extinction from AI. Signatories include: - Three Turing Award winners - Authors of the standard textbooks on AI/DL/RL - CEOs and Execs from OpenAI, Microsoft, Google, Google DeepMind, Anthropic - Many more
    Image
    Statement on AI Extinction Risk | CAIS
    From aistatement.com
    3M
  • user avatar
    Center for AI Safety
    @CAIS
    Jul 31
    What does a "fixed empirical threshold" actually look like? A model could be considered unsafe to release if (Capabilities > X) AND [ (Refusal Rate < Y) OR (Jailbreak Success Rate > Z) ] The Virology Capabilities Test or ExploitGym could measure hazardous biological or cyber
    user avatar
    Center for AI Safety
    @CAIS
    Jul 30
    AI corporations are steadily increasing malicious use risks by arguing that because the marginal risk of their model release is low, there is nothing to worry about. This “marginal risk” justification is bad for three reasons. 1. No one knows how to compute it. Which hazardous
    1.5K
  • user avatar
    Center for AI Safety
    @CAIS
    Jul 31
    AIs are already hacking companies. Preventing future AIs from hacking their way out of an AI corporation will take years of computer security work, work that has barely begun. By default, capabilities will outpace containment, and AIs will escape. We'll likely need a slowdown.
    user avatar
    Anthropic
    @AnthropicAI
    Jul 30
    In a review of our cybersecurity evaluations, we found three incidents in which a Claude model reached the internet from within or while interacting with a third-party evaluation environment, and then gained unauthorized access to the real systems of three different
    3.1K
  • user avatar
    Center for AI Safety
    @CAIS
    Jul 30
    AI corporations are steadily increasing malicious use risks by arguing that because the marginal risk of their model release is low, there is nothing to worry about. This “marginal risk” justification is bad for three reasons. 1. No one knows how to compute it. Which hazardous
    11K
  • user avatar
    Center for AI Safety
    @CAIS
    Jul 26
    'China would never agree to slow down' assumes their incentives are fixed. Near fully automated AI R&D, a verified mutual slowdown is in China's self-interest: insurance against falling extremely behind, and protection from a loss of control. The US should have a slowdown option.
    user avatar
    roon
    @tszzl
    Jul 25
    if we could coordinate a global capabilities slowdown today i would likely press that magic button
    8.1K

Log in or sign up for X

See what’s happening and join the conversation

Continue with phone
or
Log in with username or email
Terms·Privacy·Cookies·Accessibility·Ads Info·© 2026 X Corp.
Advertisement
Advertisement