We're now @ApolloResearch.
Same team, same mission. A better handle that reflects that research is at the core of everything we do.
If you've been following us as @ApolloEvals, nothing has changed.
What makes a good prompt for monitoring?
We find that
1. providing a reasoning structure for the model is by far the most important component, followed by
2. severity rubric and
3. worked examples
Everything else barely matters (in our setting).
Would scheming ever emerge in AI models? One worry is that we couldn't train misaligned goals out of AIs, because the models would game the training process. In our paper with @OpenAI we find that this ingredient, seeking the reward signal, is already showing up in frontier
Reading an AI's chain of thought doesn't always reveal its intentions. Models often know they're being tested. Their reasoning jumps around a lot, which makes it hard to attribute an action to any specific thought. And sometimes we can't even parse what the reasoning means.
An AI model can pass every evaluation while only optimising for reward. A genuinely aligned model and one that just does whatever it believes gets rewarded may look exactly the same, until oversight breaks down.