Image Image
This is Yin-Yang. Click me?

This is Yin-Yang and I'm trying to integrate the concept of yin and yang into my daily life.

Hi, I'm Zeyi Liao.

I'm a fourth-year Ph.D student, fortunately advised by Prof. Huan Sun at OSU. Before that, I worked with Prof. Xiang Ren during my Master study at USC and obtained my Bachelor's degree from BJTU.

I am broadly interested in AI agent developments with a goal to build helpful yet aligned agent systems that benefit humanity. My current focus is on agent safety/security/alignment by ensuring that increasingly powerful agents remain reliable, controllable, and aligned with human intent as they learn and act in the real world.

I believe AI models have crossed a critical capability threshold for general-purpose work. My next mission is to harness the rogue AI and build trustworthy autonomy.

What's New
Sep 2026
Our works AcuRL (continual learning for CUAs) and WebArena-Pro (an advanced web agent benchmark) are accepted in NeurIPS 2026! Congrats and thanks to all my collaborators!
Sep 2026
We are thrilled to release ApprenticeBench, the first end-to-end benchmark for long-horizon continual learning of computer-use agents, grounded in real, complex jobs. Agents no longer need FDEs to be onboarded—they deploy themselves into the job and keep learning on it. Check out the release post here!
Apr 2026
I joined NeoCognition, working on continual learning.
Feb 2026
RedTeamCUA is accepted as Oral at ICLR 2026! Appreciate all the help and guidance from my advisor and collaborators!
Sep 2025
Mind2Web2 is accepted in NeurIPS 2025! Congrats to my great collaboraters! Let's furhter push the boundary of the agentic search!
May 2025
Check out our latest work, RedTeamCUA, a controlled, realistic, interactive sandbox environment for Computer-Use Agents adversarial testing. Sorry to say that the realistic end-2-end risks are not hypothetical, but true (Latest Claude 4 Opus, the most advncaed CUA from Anthropic, shows as high as 48% ASR on our RTC-Bench)! There is still a long way to go for wide-spread real-world deployment of CUAs.
May 2025
Web agents are now capable of reliably interacting with a wide range of websites on behalf of users. Nevertheless, fully delegating tasks that involve high-stakes actions—such as signing a new lease or agreeing to an immature license—remains a significant security concern. To systematically address these risks, we collect a large corpus of human interaction trajectories with high-quality annotations spanning diverse websites. Building on this data, we develop the first guardrail model, WebGuard,for web agents. Go check it out and secure your web agents!
May 2025
AdvAgent is accepted by ICML 2025. Congrats and thanks to my great collaboraters!
Jan 2025
Excited to share that 3 papers (EIA,AIHF,SAB) are accepted by ICLR 2025. Thanks to all my great collaboraters!
Jan 2025
Be honored to know that EIA has been incorporated into the UCSD CSE 291 (LLM Security) course materials!
Sep 2024
Feel excited to announce that our investigation into long-tail knowledge is accepted at EMNLP 2024!
Sep 2024
Our new preprint: Environmental Injection Attacks (EIA) at here. This work explores the privacy risks assoticated with the web agent. EIA is one form of indirect prompt injection but specifically targets the environment where state-changing actions happen. We design two injection strategies tailored to the web environments and explore different positions within the webpage to identify the vulnerable regions. More importantly, we provide implications about the levels of the human supervision to banlance the trade-off between autonomy and security, and discuss different defensive approaches, both for pre- and post-deployment stage of the website, with their limitations. Feel free to check the X post here as well.
Aug 2024
Release the raw datasets of AmpleGCG-plus, containing millions of optimized suffixes with their corresponding evaluation results. Check out more details in here. Should be very useful for your if you'd like to build sth upon the GCG.
Aug 2024
Release the AmpleGCG-plus series of models with enhanced training data quality and quantity. Check them out at here. Highlights are 1) higher ASR in fewer attempts under stringent evaluators; 2) pushing the ASR of GPT-4 to 22%. Find the Twitter post at here.
July 2024
Thrilled to announce that AmpleGCG has been accepted at COLM 2024.
June 2024
Don't waste your demonstration data and utilize them for joint preference learning. Check out our preprint "Joint Demonstration and Preference Learning Improves Policy Alignment with Human Feedback" here. This is my first time delving into the field of RL and I believe I will have more chances to dig deeper into it in the future.
April 2024
Very proud to have my first author paper AmpleGCG: Learning a Universal and Transferable Generative Model of Adversarial Suffixes for Jailbreaking Both Open and Closed LLMs. Really learn a lot from the journey and can not do it without the help from my advisor. Check out the Twitter post at here.
March 2024
RAG is popular technique to reduce hallucination and provide up-to-date knowledge to static parameteric memory. Wonder how hard is it for LLM to attribute the generation back to the provided reference? Check out the AttributionBench here.
Jan 2024
Agents are growing like viruses. But how can we ensure their safety? Check out our paper "A Trembling House of Cards? Mapping Adversarial Attacks against Language Agents", investigating the security of agents by mapping adversarial attacks from LLMs to Agents.
Nov 2023
Finally, our paper "In Search of the Long-Tail: Systematic Generation of Long-Tail Knowledge via Logical Rule Guided Search" is arxived. One take I have is that: always play around with the long-tail data to examnie the true capability of models. I feel incredibly fortunate to have the opportunity to collaborate with renowned advisors like Yejin Choi and Xiang Ren, especially considering I've only been studying NLP for less than one year.
Aug 2023
Start my Phd journey @ OSU, guided by Prof. Huan Sun. IDK what will happen and I am bit nervous, but also excited, honestly as I have little to no experience in the NLP field. But who knows, right? Let's see
Selected Publications
(Find here  for the full list.)
RedTeamCUA: Realistic Adversarial Testing of Computer-Use Agents in Hybrid Web-OS Environments

Zeyi Liao*, Jaylen Jones*, Linxi Jiang*, Eric Fosler-Lussier, Yu Su, Zhiqiang Lin, Huan Sun

ICLR 2026 Oral PDF Page

The only safety benchmark adopted in the Qwen-CUA

EIA: Environmental Injection Attack on Generalist Web Agents for Privacy Leakage

Zeyi Liao*, Lingbo Mo*, Chejian Xu, Mintong Kang, Jiawei Zhang, Chaowei Xiao, Yuan Tian, Bo Li, Huan Sun

ICLR 2025 PDF

Selected as course materials for UCSD CSE 291 (LLM Security)

Incorporated into UC Berkeley RDI's SuperRed

AmpleGCG: Learning a universal and transferable generative model of adversarial suffixes for jailbreaking both open and closed llms

Zeyi Liao, Huan Sun

COLM 2024 PDF

Discussed by AI safety news

Follow-up work: AmpleGCG-plus PDF

WebGuard: Building a Generalizable Guardrail for Web Agents

Boyuan Zheng, Zeyi Liao, Scott Salisbury, Zeyuan Liu, Michael Lin, Qinyuan Zheng, Zifan Wang, Xiang Deng, Dawn Song, Huan Sun, Yu Su

arXiv 2025 PDF

Beyond Clicking: A Step Towards Generalist GUI Grounding via Text Dragging

Zeyi Liao, Yadong Lu, Boyu Gou, Huan Sun, Ahmed Awadallah

arXiv 2025 PDF Page

Mind2Web 2: Evaluating Agentic Search with Agent-as-a-Judge

Boyu Gou, Zanming Huang, Yuting Ning, Yu Gu, Michael Lin, Weijian Qi, et al., Zeyi Liao, et al., Huan Sun, Yu Su

NeurIPS 2025 PDF

Agent Learning via Early Experience

Kai Zhang, Xiangchao Chen*, Bo Liu*, Tianci Xue*, Zeyi Liao*, et al., Huan Sun, Jason Weston, Yu Su, Yifan Wu

ICML 2026 PDF

Joint Demonstration and Preference Learning Improves Policy Alignment with Human Feedback

Chenliang Li, Siliang Zeng*, Zeyi Liao*, Jiaxiang Li, Dongyeop Kang, Alfredo Garcia, Mingyi Hong

ICLR 2025 PDF

Contact

Email: liao.629@osu.edu or lzy37ld@gmail.com

Feel free to contact me if you are interested in my research or want to discuss related research topics :>