We're now @ApolloResearch.
Same team, same mission. A better handle that reflects that research is at the core of everything we do.
If you've been following us as @ApolloEvals, nothing has changed.
We hosted a webinar with @Tailscale on safely scaling AI coding agents.
We cover how Aperture + Watcher provide visibility, runtime controls and behavioural monitoring, with a demo of Watcher blocking data exfiltration after a prompt injection.
Video:
Apollo Research ran its first external red-teaming campaign for Anthropic’s auto mode.
Now auto mode is the default permissions mode in Claude Code. After hardening based on Apollo’s findings, the classifier's miss rate fell from 12% to 7%. 🧵
New episode with @MLStreetTalk: our research team does a deep dive on reward-seeking in AI models.
The Episode covers our recent paper with OpenAI and what these findings mean for how we evaluate frontier systems.
Filmed a week before the OpenAI/Hugging Face incident, which
What makes a good prompt for monitoring?
We find that
1. providing a reasoning structure for the model is by far the most important component, followed by
2. severity rubric and
3. worked examples
Everything else barely matters (in our setting).