About Me

I work on scaling laws and systems that make foundation-model scaling more predictable and compute-efficient. At ImageByteDance Seed, on the LLM Foundation Model team, I initiated and built the scaling-ladder framework, which uses PretrainEval proxy metrics measured on controlled small-scale training runs to (1) identify compute-optimal configurations, (2) make quantitative predictions of large-scale model behavior, and (3) expose scaling risks before committing significant compute. This work has supported decisions on model architecture and training recipes, as well as LLM interpretability and debugging, across the development of successive Seed models, including Seed1.51.5-VL1.61.82.0.

My research began in computational neuroscience before shifting to spatial understanding and, later, multimodal intelligence, with work published at ICML, NeurIPS, CVPR, ICCV, ECCV, ACL, and EMNLP. During my PhD at the ImageUniversity of Sydney, I completed research internships at ImageMeta FAIR, ImageAWS AI Lab, ImageMicrosoft GenAI, and ImageGoogle AI. I also earned my bachelor's degree there with First Class Honours and the University Medal.

Email: chaoyivision@gmail.com

Research

Scaling Law for Quantization-Aware Training thumbnail
Mengzhao Chen, Chaoyi Zhang, Jing Liu, et al.
Scaling Law Established empirical scaling relationships for quantization-aware training across model scale and numerical precision.
INT v.s. FP: A Comprehensive Study of Fine-Grained Low-bit Quantization Formats thumbnail
Mengzhao Chen, Meng Wu, Hui Jin, …, Chaoyi Zhang, et al.
Pretraining Systematically compared fine-grained integer and floating-point formats for low-bit model training and inference.
Model Merging in Pre-training of Large Language Models thumbnail
Yunshui Li, Yiyuan Ma, Shen Yan, Chaoyi Zhang, et al.
Pretraining Investigated model merging as a mechanism within pretraining rather than only as a post-training technique.
Virtual Width Networks thumbnail
ByteDance Seed
Architecture Explored virtual network width as a controllable axis for studying and improving large-model training.

* A full list of my publications can be found in my Google Scholar.