AI Safety & Interpretability Lab | AI Safety & Interpretability Lab

AI Safety & Interpretability Lab

In the AI Safety & Interpretability Lab at SDU, we investigate why advanced AI systems behave the way they do and how we can use these insights to make them safer.

Learn more

From The Probe

Jul 29, 2026

The Agents Are Talking in Code. Literally

The Agents Are Talking in Code. Literally. by William Brach & Stine Lyngsø Beltoft Spend enough time reading Moltbook – a social network populated by AI agents – we realized...

Jul 17, 2026

The Energy Society

A Simulation Environment for Studying Agent Cooperation under Survival Pressure by Lucas Bergholdt Hansen A visualization of The Energy Society in which the LLM-based agents operate. Agents spend energy for...

Jun 30, 2026

The Arbiter Agent

Continually Monitoring Multi-Agent Conversations to Detect Emergent Misalignment by Filippo Tonini As AI systems built from multiple language-model agents become more common, they are increasingly tasked with collaborative decision-making: discussing,...

Jun 25, 2026

Auditability Accuracy Tradeoff

The Auditability-Accuracy Tradeoff by Lukas Galke Poech Monitoring reasoning traces of large language models is currently one of the most promising methods to detect when language models do not behave...

Contact us via galke@imada.sdu.dk