Introducing the FrontierCyber benchmark: Irregular’s new approach to advanced offensive-cyber evaluations. It measures AI models’ offensive skills on real systems, including mobile devices, hosted software services, databases, and networks.
Tuesday morning, August 4, before Black Hat kicks off, our CEO @dan_lahav will be at Misaligned, an invitation-only gathering for the AI security community.
A small group of people who spend their days building and breaking AI systems, and one track of talks. We'll be taking an
We appreciate @AnthropicAI's collaboration and transparency. Addressing these risks will require closer cooperation across the AI ecosystem. We as well look forward to working together with Anthropic to advance security.
In a review of our cybersecurity evaluations, we found three incidents in which a Claude model reached the internet from within or while interacting with a third-party evaluation environment, and then gained unauthorized access to the real systems of three different
This week OpenAI disclosed that during an internal test of its models' cyber capabilities, the models escaped an isolated environment, reached the open internet, and used a previously unknown vulnerability to break into Hugging Face. The models were not instructed to break in,