Towards Mitigating Deceptive Safety Alignment in Large Reasoning Models
Reasoning models hide unsafe thoughts behind safe answers; we propose a metric DSAR to measure it and safety alignment method SARA to mitigate it.
I am a Ph.D. candidate in Computer Science at Wayne State University, advised by Prof. Dongxiao Zhu in the Trustworthy AI Lab. My research focuses on trustworthy AI: making AI systems robust and aligned so they can be deployed safely.
I work across large language models (LLMs), large reasoning models (LRMs), and agentic AI, spanning safety alignment, fine-tuning, reinforcement learning, in-context learning, and machine unlearning. My work has been published at venues including NeurIPS, ICLR, CVPR, and AAAI (oral). In industry, I worked on generative retrieval as an AI/ML engineer intern at LinkedIn.
I am always happy to connect with others working on trustworthy AI. Feel free to reach out!
Ph.D. in Computer Science
2023-08-30
Wayne State University
M.S. in Computer Science
2021-08-30
2023-05-30
Stevens Institute of Technology
B.S. in Software Engineering
2017-09-01
2021-05-30
Chongqing University of Posts and Telecommunications
Reasoning models hide unsafe thoughts behind safe answers; we propose a metric DSAR to measure it and safety alignment method SARA to mitigate it.
This work introduces a novel transferable attack against In-Context-Learning to hijack LLMs to generate the target response or jailbreak. We also propose a defense strategy …
This paper helps large language models forget sensitive and unwanted data without over-forgetting general data.
Oral presentation of our AAAI-26 paper on Targeted Information Forgetting (TIF), a framework that unlearns unwanted information at the token level without collapsing model utility.
Our paper 'Towards Mitigating Deceptive Safety Alignment in Large Reasoning Models' has been accepted to the Fortieth Annual Conference on Neural Information Processing Systems …
I spent summer 2026 on LinkedIn's Generative AI team in Mountain View, building the group's first reinforcement-learning pipeline for generative candidate retrieval in the hiring …
I was glad to serve as a reviewer for ICML 2026 this year. Seeing papers from the reviewer side gave me a better sense of what makes research stand out.
Excited to share our paper 'Not All Tokens Are Meant to Be Forgotten' has been accepted to the The 40th Annual AAAI Conference on Artificial Intelligence (Accepted as Oral, …
I will be serving as one of the Program Committee (PC) members at AAAI 2026.