I work on scaling laws and systems that make foundation-model scaling more predictable and compute-efficient. At ByteDance Seed, on the LLM Foundation Model team, I initiated and built the scaling-ladder framework, which uses PretrainEval proxy metrics measured on controlled small-scale training runs to (1) identify compute-optimal configurations, (2) make quantitative predictions of large-scale model behavior, and (3) expose scaling risks before committing significant compute. This work has supported decisions on model architecture and training recipes, as well as LLM interpretability and debugging, across the development of successive Seed models, including Seed1.51.5-VL1.61.82.0.
My research began in computational neuroscience before shifting to spatial understanding and, later, multimodal intelligence, with work published at ICML, NeurIPS, CVPR, ICCV, ECCV, ACL, and EMNLP. During my PhD at the University of Sydney, I completed research internships at Meta FAIR, AWS AI Lab, Microsoft GenAI, and Google AI. I also earned my bachelor's degree there with First Class Honours and the University Medal.
Agentic MLLMIntroduced adaptive perception magnification to let vision-language models inspect visual evidence during decoding and reduce hallucination.