I’m interested in how frontier models take part in constructing their successors, through environments, algorithms, infrastructure, and harnesses. I view it as a horizontal endeavor, one I hope to realize through the right way of scaling.
I am a second-year Ph.D. student in Machine Learning at Georgia Tech, co-advised by Bo Dai and Chao Zhang.
Currently, I am a Student Researcher at Google DeepMind in Mountain View, where I work on agentic MLE for Gemini. I contribute to the Gemini 3.5 and 3.6 Flash models.
Before Georgia Tech, I received my B.Eng. in Automation from Tsinghua University.
I welcome conversations about research, collaboration, and opportunities: reach me at rqiang6@gatech.edu.
Research
My research aims at self-improving AI, organized as a stack in which every layer has to hold:
- Data & Environments. Interactive playgrounds where agents run the real experiment loop, and automated pipelines that manufacture verifiable tasks at scale (MLE-Dojo, MLE-Smith).
- Algorithms. Post-training that stays dense, reliable, and on-policy over long horizons (Agent DAgger).
- Infrastructure. Modular systems that schedule and scale agentic reinforcement learning (STACX).
- Harnesses. Hierarchical orchestration that sustains long-horizon optimization and research (Matryoshka Agent).
I wish to find the right way of scaling for Recursive Self-Improvement.
I strive for research and projects that are scalable and have real impact on frontier models.
I contribute to frontier models such as Gemini (3.5 Flash, 3.6 Flash).
MLE-Dojo and MLE-Smith have been widely adopted for building MLE/RSI agents and systems, and have proven scalable (e.g., OpenRSI).
Publications & Preprints
- ICLR 2026 MLE-Smith: Scaling MLE Tasks with Automated Multi-Agent Pipeline
- NeurIPS 2025 · Datasets & Benchmarks MLE-Dojo: Interactive Environments for Empowering LLM Agents in Machine Learning Engineering
- NeurIPS 2025 Matryoshka Pilot: Learning to Drive Black-Box LLMs with LLMs
- COLM 2025 Language Model Uncertainty Quantification with Attention Chain
- NeurIPS 2024 HYDRA: Model Factorization Framework for Black-Box LLM Personalization
- NAACL 2024 AutoLoRA: Automatically Tuning Matrix Ranks in Low-Rank Adaptation Based on Meta Learning
- arXiv · 2026 Matryoshka Agent: Unfolding Sub-Agents for Long-Horizon Machine Learning Engineering
- arXiv · 2026 Revisiting DAgger in the Era of LLM-Agents
- arXiv · 2026 Forward-Free Diffusion Language Models
- arXiv · 2025 Towards Better Instruction Following Retrieval Models
* equal contribution · † equal second authorship
Software
Experience
- Google DeepMind, Mountain View Student Researcher · agentic MLE for Gemini Feb 2026 – present
Education
- Georgia Institute of Technology, Atlanta Ph.D. student in Machine Learning · co-advised by Bo Dai and Chao Zhang Aug 2024 – present
- Tsinghua University, Beijing B.Eng. in Automation Sep 2020 – Jun 2024
Academic Service
Reviewer for NeurIPS, ICLR, ICML, and ACL Rolling Review.