intern @expsecai | eu/acc | msc data science | ai safety & alignment | long horizon, instrumental convergence, multi-agent failure modes | views are my own
Lots of discussion out there about our next model(!), so I wanted to give an early look as soon as possible. Introducing Gemini 4 Argon!
It shows frontier performance in complex workflows, cyber defense and software engineering. Teams are using it extensively at Google, from
Our original attack allows extracting reasoning of the recent frontier models, including Astra and Sol 6.1.
On reasoning effort MAX both Astra and Sol become very aware of their "token budgets", and eventually start saving tokens by omitting white spaces
More examples on
Two months after our reasoning extraction attack release, we audited mitigations introduced since then and found that we still can extract reasoning from Astra/Sol-6.1 via third-party API providers.
We disclosed our findings to OpenAI and Anthropic...
openai.com/index/disrupti…
if you ask different models where they would leave a message under certain constraints they all give very similar answers and name the same sites.
i asked multiple models the same quesition which i derived based on the constraints the models had during their web-retrieval eval.