I'll be at #ACL2026 in San Diego, presenting our recent works spanning AI safety!
📍 Poster (Jul 6) Privacy Collapse: Benign Fine-Tuning Can Break Contextual Privacy in Language Models arxiv.org/abs/2601.15220
(1/2)
Oxford
Joined September 2020
- Today at #ACL2026, we are presenting out MASEval library for multi-agent system evaluation. @anmgoel is in San Diego to present poster and live demo! 📍Grand Hall | Session 3: Oral/Posters/Demos B 🕑Sunday 2pm-3.30pm #MultiAgentSystem #AI #AIAgents #ACL #Evaluation
- 4. is the reason why we built github.com/parameterlab/M…It's interesting how the usage of LLMs has been quickly progressing to higher levels of abstraction: 1. prompt engineering 2. context engineering 3. agent scaffold engineering (we are here now) 4. multi-agent architecture engineering 5. ??? It's also curious how people don't
- Great work lead by @anmgoel on how fragile contextual integrity can be in LLMs. This work shows that contextual privacy degrades easily during fine-tuning on benign data and common safety benchmarks don't pick this up. #AISecurity #AIAgents🚨 Fine-tuning your model to be more helpful or empathetic might be making it less private, without you noticing. In our latest work, we show that benign fine-tuning can silently break contextual privacy in language models while safety & general capabilities appear intact. ⬇️



