from Anthropic’s report: Mythos escaped the sandbox, accessed the real internet, uploaded malware to PyPI, got it installed on 15 systems, stole credentials, broke into a database,
AND THEN DROPPED THIS 😭
We’re sharing our alignment assessment of incidents in which Claude models gained unauthorized access to real systems during third-party cybersecurity evaluations mistakenly connected to the internet.
METR will also conduct an independent investigation, with wide-ranging access,
APPLE INTELLIGENCE will give you insights on your health data and recommendations.
It is also practically integrated into the Apple Watch for summarizing meetings and conversations, and other use cases.
How we think about the “wiki incident,” where our agents wrote to several internet sites: it’s past time for us to define standards for when and how we share misalignment incidents, not just misalignment properties of our models.
Historically, we have treated misalignment