About Me
Hi! I’m Zoey. I work at Amazon SFAI developing the language model behind the Rufus assistant. I graduated as a Computer Science PhD from UIUC in 2024 advised by Prof. Jiawei Han. I’m very fortunate to be part of two generous families: the Data Mining Group and the BLENDER lab.
These days I work on RL post-training for language models, from the algorithms to the training infrastructure behind them. In 2024–2025 I focused on data selection and scaling laws for pretraining. During my PhD I built semi-supervised and weakly supervised methods for information extraction (pre-GPT), then moved to knowledge editing and hallucination mitigation.
Recent Writing
Latest
Revisiting the Predictability of RLVR
RELEX and AlphaRL find that RLVR weight updates are low rank and evolve near-linearly, and conclude that RLVR training is predictable. Their low-rank measurements reproduce. But a random walk produces the same measurements, rank-1 recovers the gain only when it is fitted to the endpoint it reconstructs, and on our runs extrapolation fails even at RELEX’s own horizon.
Series
Off-policy corrections
Where the policy that samples rollouts and the policy being trained come apart in LLM RL, and when a correction is worth it.
