About Me

Hi! I’m Zoey. I work at Amazon SFAI developing the language model behind the Rufus assistant. I graduated as a Computer Science PhD from UIUC in 2024 advised by Prof. Jiawei Han. I’m very fortunate to be part of two generous families: the Data Mining Group and the BLENDER lab.

These days I work on RL post-training for language models, from the algorithms to the training infrastructure behind them. In 2024–2025 I focused on data selection and scaling laws for pretraining. During my PhD I built semi-supervised and weakly supervised methods for information extraction (pre-GPT), then moved to knowledge editing and hallucination mitigation.

[Google Scholar] · [GitHub] · [Twitter]

Recent Writing

Latest

Revisiting the Predictability of RLVR

33 min read , In progress

RELEX and AlphaRL find that RLVR weight updates are low rank and evolve near-linearly, and conclude that RLVR training is predictable. Their low-rank measurements reproduce. But a random walk produces the same measurements, rank-1 recovers the gain only when it is fitted to the endpoint it reconstructs, and on our runs extrapolation fails even at RELEX’s own horizon.

Series

Off-policy corrections

Where the policy that samples rollouts and the policy being trained come apart in LLM RL, and when a correction is worth it.

  1. Off-Policy Corrections in LLM RL Training Mar 1, 2026 · 30 min read
  2. The Infrastructure Cost of MoE Routing Replay Apr 26, 2026 · 16 min read
  3. A Field Guide to Training–Inference Corrections Sep 5, 2026 · 17 min read
  4. Signal or Noise? An SNR Criterion for Trusting Your Importance Ratio Sep 7, 2026 · 38 min read
All posts →