1. X
  2. Anmol Goel
Log inSign up
Anmol Goel
448 posts
Image
user avatar
Anmol Goel
@anmgoel
AI Safety x Privacy @ELLISForEurope PhD @UKPLab, @TUDarmstadt and @UCPH_Research | Prev @iiit_hyderabad |
Darmstadt 🇩🇪
anmol.goel.ai
Joined September 2019
1,571
Following
667
Followers
RepliesRepliesMediaMedia
  • Pinned
    user avatar
    Anmol Goel
    @anmgoel
    Jul 17
    Do computer-use agents (CUAs) share info appropriately in multi-app personal workflows? We release AgentCIBench: an evaluation harness to test this. Project: ukplab.github.io/arxiv2026-agen… 🚨An agent that aces your office task will also happily share your medical appointments.
    Image
    00:00
  • user avatar
    Anmol Goel
    @anmgoel
    Jul 13
    This is amazing!!
    user avatar
    Zane Chee
    @injaneity
    Jul 12
    pi-computer-use v0.4.3 is here! an update to bring us to parity with codex (who to their credit, stepped up their game massively in their last update): - a ghost cursor now shows what computer use is working on - parallel tasks allow subagents to coordinate across different apps
    Image
    00:00
  • user avatar
    Anmol Goel
    @anmgoel
    Jun 29
    I'll be at #ACL2026 in San Diego, presenting our recent works spanning AI safety! 📍 Poster (Jul 6) Privacy Collapse: Benign Fine-Tuning Can Break Contextual Privacy in Language Models arxiv.org/abs/2601.15220 (1/2)
    Image
    00:00
  • user avatar
    Anmol Goel
    @anmgoel
    Apr 7
    Accepted to ACL 2026 Main Conference🇺🇸! We show benign fine-tuning on tasks like empathy, and code debugging, systematically weakens contextual privacy while preserving general performance, revealing a key gap in current safety evaluations. Paper: arxiv.org/abs/2601.15220
    user avatar
    Anmol Goel
    @anmgoel
    Feb 3
    🚨 Fine-tuning your model to be more helpful or empathetic might be making it less private, without you noticing. In our latest work, we show that benign fine-tuning can silently break contextual privacy in language models while safety & general capabilities appear intact. ⬇️
    Image
  • user avatar
    Anmol Goel
    @anmgoel
    Mar 23
    Multi-agent systems are about more than just the underlying models; the entire agentic harness needs to be evaluated. We built MASEval: a framework designed to handle the full evaluation lifecycle for you. Check it out here! 👇 github.com/parameterlab/M…
    user avatar
    Cornelius Emde
    @CorEmde
    Mar 23
    1/ Evaluating a single agent harness is hard. Evaluating a multi-agent system? That's a whole different problem. Most eval tools treat the model as the unit of analysis. But in multi-agent systems, the system is what matters. That's why we built MASEval 🧵 #Agents #AI #Eval
    Image

Log in or sign up for X

See what’s happening and join the conversation

Continue with phone
or
Log in with username or email
Terms·Privacy·Cookies·Accessibility·Ads Info·© 2026 X Corp.
Advertisement
Advertisement