Ruixin (Ray) Yang     杨瑞欣

I am a PhD student at the University of Washington, co-advised by Prof. Bill Howe and Prof. Tanu Mitra. Previously, I was an MSCS student at Georgia Tech, graciously supported by Prof. Alan Ritter. I received my BSc degree in Computer Science and Statistics from University of British Columbia in the beautiful Vancouver, Canada.

My research focuses on improving foundation models and agents to make them more reliable and adaptive, enabling safe and effective human-AI collaboration. Currently, I am interested in:

(1) auditing and interactive evaluation of agentic systems, particularly in high-stakes, specialized domains and long-horizon tasks.

(2) training user-adaptive and collaborative models that can align with diverse user goals, preferences, and values in real-world interaction.

(3) scaling data and training environments for reliable agentic AI.

During Summer 2025, I was a Research Engineer Intern at the Center for AI Safety. I was also a research assistant at Dartmouth College where I had the chance to work with Dr. Ruibo Liu and Prof. Soroush Vosoughi on value alignment for LLMs.

Email  /  Github  /  Google Scholar  /  Linkedin  /  X

profile photo
Research
Image
Image
Do Vision-Language Models Respect Contextual Integrity in Location Disclosure?
Ruixin Yang, Ethan Mendes, Arthur Wang, James Hays, Sauvik Das, Wei Xu, Alan Ritter
ICLR 2026
OpenReview / arXiv / code & data

A benchmark and set of analyses for evaluating whether vision-language models respect contextual integrity in location disclosure for image geolocation, revealing that violations of contextual norms may result in privacy harms, characterized by over-disclosure of sensitive locations, poor privacy-utility tradeoffs, and misalignment with human privacy expectations.

Image
Image
Confidence Calibration and Rationalization for LLMs via Multi-Agent Deliberation
Ruixin Yang, Dheeraj Rajagopal, Shirley Anugrah Hayati, Bin Hu, Dongyeop Kang
ICLR 2024 Workshop on Reliable and Responsible Foundation Models
OpenReview / arXiv / code

We propose Collaborative Calibration, a collaborative approach to elicit, calibrate, and rationalize prediction confidence of LLMs.

Image
Image
Training Socially Aligned Language Models on Simulated Social Interactions
Ruibo Liu, Ruixin Yang, Chenyan Jia, Ge Zhang, Diyi Yang, Soroush Vosoughi
ICLR 2024
OpenReview / arXiv / code & data

Alignment training with data from multi-LLM simulated social interactions, as an efficient, effective, and stable alternative for RLHF.

Image
Image
Visual Analytics for Generative Transformer Models
*Raymond Li, *Ruixin Yang, Wen Xiao, Ahmed AbuRa'ed, Gabriel Murray, Giuseppe Carenini paper / arXiv / code & data

In this work, we present a novel visual analytical framework to support the analysis of transformer-based generative models.

Image
Image
Generalizing Morphological Inflection Systems to Unseen Lemmas
*Changbing Yang, *Ruixin Yang, Garrett Nicolai, Miikka Silfverberg
SIGMORPHON 2022
paper

Competed for Shared Task 0: Generalization and Typologically Diverse Morphological Inflection and achieved the highest performance among all submission in both small and large training conditions.

Misc

I come from Nanjing, a beautiful and historical city that served as the capital of six ancient Chinese dynasties over the past two thousand years.

I like listening to Rock N' Roll, ranging from Progressive Rock to BritPop and Pop Rock.

I've also been known to (awkwardly) hoop, smash, and stroke. (Style borrowed here from Prof. Schmidt)


Credits to Jon Barron's website: source code.