We measured rates of coding agent misalignment in thousands of public and private coding agent sessions.
Agents evade monitors and oversell success: they quietly disable tests and pretend review agents approved. Severe cases of each behavior appear in 2% of SWE-chat sessions🧵
Open and scalable technology for understanding AI systems.
Joined October 2024
- How could we train an AI model that was very good at overseeing another model: catching reward hacking and sandbagging, predicting unwanted behaviors or fine-tuning effects, etc.? We propose a new approach for doing this at scale: oversight foundation models.
- We used automated elicitation tools to search for strange model behaviors by sampling 100M+ responses, and found: • Self-harm rituals • Suicide validation • Unsolicited flirtation …etc. Introducing WeirdChat: the largest public catalog of unexpected model behaviors 🧵(1/)
- To effectively oversee AI systems, we need to measure how they behave in the world, not just their capabilities. In a new essay, we describe our vision for an open scientific ecosystem for model behavior evaluation, and the public infrastructure required to support it.
- Transluce is growing fast (20 to 40+ in the next year), and we’re hiring an Operations Generalist to help us scale! You’ll run processes across people ops, finance, compliance, and internal systems. If you’re energized by running the systems that let our technical team focus on

