Log inSign up
Center for AI Safety
322 posts
Center for AI Safety profile banner
@CAIS

Center for AI Safety

@CAIS
Reducing societal-scale risks from AI.
San Francisco
safe.ai
Joined August 2022
3
Following
10.8K
Followers
RepliesRepliesRepostsRepostsMediaMediaArticlesArticles

Log in or sign up for X

See what’s happening and join the conversation

Continue with phone
or
Log in with username or email
Terms·Privacy·Cookies·Accessibility·Ads Info·© 2026 X Corp.
  • Pinned
    @CAIS
    Center for AI Safety
    @CAIS
    May 30, 2023
    We’ve released a statement on the risk of extinction from AI. Signatories include: - Three Turing Award winners - Authors of the standard textbooks on AI/DL/RL - CEOs and Execs from OpenAI, Microsoft, Google, Google DeepMind, Anthropic - Many more
    Image
    Statement on AI Extinction Risk | CAIS
    From aistatement.com
    149
  • @CAIS
    Center for AI Safety
    @CAIS
    Aug 14
    > "We will release the weights in two weeks... once safety evaluation and hardening are complete." Hardening society against AI cyberattacks will take more than two weeks. Defending against AI cyberattacks would require upgrading critical infrastructure and other computers so
    @Zai_org
    Z.ai
    @Zai_org
    Aug 14
    Introducing GLM-5.3: Built to Code. Ready for Cyber Defense. - Top-tier coding and agentic capabilities, achieved through post-training on the 743B base model - A major leap in cybersecurity, setting a new standard among open models Tech Blog: z.ai/blog/glm-5.3
    Image
    6
  • @CAIS
    Center for AI Safety
    @CAIS
    Aug 14
    Relying on an AI to "think out loud" (called chain-of-thought monitoring) is not a long-term safety solution. 1. As AI models get larger, more thinking happens deep within their layers before they utter a word. 2. AIs can already alter their chain of thought when prompted, so
    @kotekjedi_ml
    Alexander Panfilov
    @kotekjedi_ml
    Aug 11
    Replying to @kotekjedi_ml
    2) Illegible reasoning: We confirm prior reports by @ApolloResearch: OpenAI models sometimes reason in alien-like language, referring to themselves as “we” or “it,” or spiraling into cursed loops of “vantages,” “marinades,” and “watchers.” CoT-monitoring people are doing God’s
    Image
    Image
    8
  • @CAIS
    Center for AI Safety
    @CAIS
    Jul 31
    What does a "fixed empirical threshold" actually look like? A model could be considered unsafe to release if (Capabilities > X) AND [ (Refusal Rate < Y) OR (Jailbreak Success Rate > Z) ] The Virology Capabilities Test or ExploitGym could measure hazardous biological or cyber
    @CAIS
    Center for AI Safety
    @CAIS
    Jul 30
    AI corporations are steadily increasing malicious use risks by arguing that because the marginal risk of their model release is low, there is nothing to worry about. This “marginal risk” justification is bad for three reasons. 1. No one knows how to compute it. Which hazardous
    3
  • @CAIS
    Center for AI Safety
    @CAIS
    Jul 31
    AIs are already hacking companies. Preventing future AIs from hacking their way out of an AI corporation will take years of computer security work, work that has barely begun. By default, capabilities will outpace containment, and AIs will escape. We'll likely need a slowdown.
    @AnthropicAI
    Anthropic
    @AnthropicAI
    Jul 30
    In a review of our cybersecurity evaluations, we found three incidents in which a Claude model reached the internet from within or while interacting with a third-party evaluation environment, and then gained unauthorized access to the real systems of three different
    10
Advertisement
Advertisement