Xiangyu Zhou 👨‍💻

Xiangyu Zhou

(he/him)

Ph.D. Candidate in Computer Science

About Me

I am a Ph.D. candidate in Computer Science at Wayne State University, advised by Prof. Dongxiao Zhu in the Trustworthy AI Lab. My research focuses on trustworthy AI: making AI systems robust and aligned so they can be deployed safely.

I work across large language models (LLMs), large reasoning models (LRMs), and agentic AI, spanning safety alignment, fine-tuning, reinforcement learning, in-context learning, and machine unlearning. My work has been published at venues including NeurIPS, ICLR, CVPR, and AAAI (oral). In industry, I worked on generative retrieval as an AI/ML engineer intern at LinkedIn.

I am always happy to connect with others working on trustworthy AI. Feel free to reach out!

Education

Ph.D. in Computer Science

2023-08-30

Wayne State University

M.S. in Computer Science

2021-08-30
2023-05-30

Stevens Institute of Technology

B.S. in Software Engineering

2017-09-01
2021-05-30

Chongqing University of Posts and Telecommunications

Interests

Trustworthy AI Large Language Models Large Reasoning Models Generative Retrieval Agentic AI
Featured Publications
Towards Mitigating Deceptive Safety Alignment in Large Reasoning Models featured image

Towards Mitigating Deceptive Safety Alignment in Large Reasoning Models

Reasoning models hide unsafe thoughts behind safe answers; we propose a metric DSAR to measure it and safety alignment method SARA to mitigate it.

avatar
Xiangyu Zhou
•
Hijacking Large Language Models via Adversarial In-Context Learning featured image

Hijacking Large Language Models via Adversarial In-Context Learning

This work introduces a novel transferable attack against In-Context-Learning to hijack LLMs to generate the target response or jailbreak. We also propose a defense strategy …

avatar
Xiangyu Zhou
•
Not all tokens are meant to be forgotten featured image

Not all tokens are meant to be forgotten

This paper helps large language models forget sensitive and unwanted data without over-forgetting general data.

avatar
Xiangyu Zhou
•
Recent Publications
Recent & Upcoming Talks
Not All Tokens Are Meant to Be Forgotten featured image

Not All Tokens Are Meant to Be Forgotten

Oral presentation of our AAAI-26 paper on Targeted Information Forgetting (TIF), a framework that unlearns unwanted information at the token level without collapsing model utility.

avatar
Xiangyu Zhou
•
Recent News
🎉 Paper accepted by NeurIPS-26 featured image

🎉 Paper accepted by NeurIPS-26

Our paper 'Towards Mitigating Deceptive Safety Alignment in Large Reasoning Models' has been accepted to the Fortieth Annual Conference on Neural Information Processing Systems …

avatar
Xiangyu Zhou
•
💼 Summer 2026 at LinkedIn as an AI/ML Engineer Intern featured image

💼 Summer 2026 at LinkedIn as an AI/ML Engineer Intern

I spent summer 2026 on LinkedIn's Generative AI team in Mountain View, building the group's first reinforcement-learning pipeline for generative candidate retrieval in the hiring …

avatar
Xiangyu Zhou
•
Will be serving as one of the reviewer at ICML 2026 featured image

Will be serving as one of the reviewer at ICML 2026

I was glad to serve as a reviewer for ICML 2026 this year. Seeing papers from the reviewer side gave me a better sense of what makes research stand out.

avatar
Xiangyu Zhou
•
🎉Paper accepted by AAAI-26 featured image

🎉Paper accepted by AAAI-26

Excited to share our paper 'Not All Tokens Are Meant to Be Forgotten' has been accepted to the The 40th Annual AAAI Conference on Artificial Intelligence (Accepted as Oral, …

avatar
Xiangyu Zhou
•
Serving as one of the Program Committee (PC) members at AAAI 2026 featured image

Serving as one of the Program Committee (PC) members at AAAI 2026

I will be serving as one of the Program Committee (PC) members at AAAI 2026.

avatar
Xiangyu Zhou
•