🚨 Our paper SkillSafetyBench is accepted to #EMNLP2026 Main!
Agent skills are becoming the plugin ecosystem of LLM agents. We ask a simple question: what if the attacker is not the user, but the skill itself?
📄 arxiv.org/abs/2605.12015
🚨False Sense of Security: Our new paper identifies a critical limitation in representation probing-based malicious input detection—purported "high detection accuracy" may confer a false sense of security:
arxiv.org/pdf/2509.03888
We will present our #ICML2024 poster tomorrow at Hall C 4-9 #2411 🥳 Welcome to have a look
Paper: On the Duality Between Sharpness-Aware Minimization and Adversarial Training
PDF: arxiv.org/pdf/2402.15152
Code: github.com/weizeming/SAM_…
🧵(1/8) I've completed my studies at UC Berkeley and resumed my undergraduate studies at Peking University 😆
✨ Some highlights from my visit to Berkeley include:
Update: I’ve just started a visiting student program at UC Berkeley for 4 months.
Please DM me if you’re also here. Looking forward to meeting old and new friends!