1. X
  2. Cornelius Emde
Log inSign up
Cornelius Emde
76 posts
Image
user avatar
Cornelius Emde
@CorEmde
AI Security | AI Agents | ML Robustness | PhD @UniofOxford and @OxfordTVG | ex RS @Wise
Oxford
Joined September 2020
678
Following
167
Followers
RepliesRepliesMediaMedia
  • user avatar
    Cornelius Emde
    @CorEmde
    Jul 5
    Check out our work on how fine-tuning breaks contextual privacy in LLMs. Very important for the security of fine-tuned agents! Monday 9am Great Hall at #ACL2026 in San Diego. Lead by @anmgoel. #LLM #AI #AISafety
    user avatar
    Anmol Goel
    @anmgoel
    Jun 29
    I'll be at #ACL2026 in San Diego, presenting our recent works spanning AI safety! 📍 Poster (Jul 6) Privacy Collapse: Benign Fine-Tuning Can Break Contextual Privacy in Language Models arxiv.org/abs/2601.15220 (1/2)
    Image
    00:00
  • user avatar
    Cornelius Emde
    @CorEmde
    Jul 5
    Today at #ACL2026, we are presenting out MASEval library for multi-agent system evaluation. @anmgoel is in San Diego to present poster and live demo! 📍Grand Hall | Session 3: Oral/Posters/Demos B 🕑Sunday 2pm-3.30pm #MultiAgentSystem #AI #AIAgents #ACL #Evaluation
    user avatar
    Cornelius Emde
    @CorEmde
    Mar 23
    1/ Evaluating a single agent harness is hard. Evaluating a multi-agent system? That's a whole different problem. Most eval tools treat the model as the unit of analysis. But in multi-agent systems, the system is what matters. That's why we built MASEval 🧵 #Agents #AI #Eval
    Image
  • user avatar
    Cornelius Emde
    @CorEmde
    Mar 24
    4. is the reason why we built github.com/parameterlab/M…
    user avatar
    Maksym Andriushchenko
    @maksym_andr
    Mar 24
    It's interesting how the usage of LLMs has been quickly progressing to higher levels of abstraction: 1. prompt engineering 2. context engineering 3. agent scaffold engineering (we are here now) 4. multi-agent architecture engineering 5. ??? It's also curious how people don't
  • user avatar
    Cornelius Emde
    @CorEmde
    Mar 23
    1/ Evaluating a single agent harness is hard. Evaluating a multi-agent system? That's a whole different problem. Most eval tools treat the model as the unit of analysis. But in multi-agent systems, the system is what matters. That's why we built MASEval 🧵 #Agents #AI #Eval
    Image
  • user avatar
    Cornelius Emde
    @CorEmde
    Mar 18
    Great work lead by @anmgoel on how fragile contextual integrity can be in LLMs. This work shows that contextual privacy degrades easily during fine-tuning on benign data and common safety benchmarks don't pick this up. #AISecurity #AIAgents
    user avatar
    Anmol Goel
    @anmgoel
    Feb 3
    🚨 Fine-tuning your model to be more helpful or empathetic might be making it less private, without you noticing. In our latest work, we show that benign fine-tuning can silently break contextual privacy in language models while safety & general capabilities appear intact. ⬇️
    Image

Log in or sign up for X

See what’s happening and join the conversation

Continue with phone
or
Log in with username or email
Terms·Privacy·Cookies·Accessibility·Ads Info·© 2026 X Corp.
Advertisement
Advertisement