About Me
I develop efficient foundation models and agents that move from language to physical intelligence. My research spans natural language processing, foundation-model training, efficient model architecture and inference, and embodied AI, with a growing focus on VLA models, robotic agents, and hardware–model co-design for real-world deployment.
A central question across my work is how to build capable AI systems under real constraints on compute, memory, latency, energy, and interaction: from language models running on edge devices to embodied agents acting on robots and vehicles.
Currently, I am a Research Fellow at the Bayes Centre, University of Edinburgh, working in collaboration with Prof. Luo Mai, Prof. Jeff Pan, and Prof. Jun Wang. I also serve as a Visiting Research Fellow at Li Auto. Prior to joining the University of Edinburgh, I was a Research Assistant at the Hong Kong University of Science and Technology (Guangzhou), where I worked with Prof. Lei Chen and Prof. Lionel M. Ni. I obtained my Ph.D. at Shanghai Jiao Tong University, where I was fortunate to be supervised by Prof. Weinan Zhang, Prof. Luoyi Fu, and Prof. Xinbing Wang. During my early research career, I interned with the Data Team at TikTok and worked as an Applied Scientist Intern at Amazon Shanghai AI Lab. In 2021, I was selected for the Wenjun Wu Honored Ph.D. Class.
Research
My current research is organised around three connected directions:
- Efficient Physical AI & Robotics. VLA and world-action models, embodied agents, and hybrid learned/classical robot systems, with deployments across robotic manipulation, mobile robots, and intelligent vehicles.
- Efficient Foundation Models & Agents. Hardware-aware model architecture, scaling laws, efficient attention and inference, and long-context agent systems for resource-constrained deployment.
- NLP, Foundation Models & AI for Science. Pre-training, post-training, reasoning, agents, and structured knowledge for language and scientific domains.
My long-term goal is to understand how models, agents, robot skills, and hardware should be co-designed so that increasingly capable AI can operate reliably in the physical world.
News
- [2026-09] I am honored to serve as an Area Chair for ICLR 2027!
- [2026-09] Dancing in Fetters, our hardware co-design scaling law for on-device LLMs, is accepted by NeurIPS 2026!
- [2026-09] RooflineBench is accepted by DAI 2026!
- [2026-09] EffVLA is accepted by CoRL 2026!
- [2026-08] Invited talk on “Large Discovery Model and Geoscience” at the 3rd International Symposium on Deep Underground Science and Engineering in Glasgow!
- [2026-04] 2 papers accepted by ACL Main Conference 2026, 2 papers accepted by ICLR 2026 (incl. SpatialViz-Bench), and ContextPilot accepted by MLSys 2026!
- [2026-04] I hosted an academic workshop at the Bayes Centre, University of Edinburgh, titled “Edge AI Agent Workshop”, where I also delivered a talk on “Efficient AI Agent on the Edge”.
Selected Research Highlights
🤖 Efficient Physical AI & Robotics
Efficient VLA CoRL 2026
A controlled, latency-aware study of modular VLA design. We identify where additional model capacity actually pays off, derive an efficient VLA recipe, and validate transfer from simulation to real robotic manipulation.
arXiv
PhysicalAgent Banbu-supported
An embodied-agent testbed integrating VLA policies, SLAM/navigation, multimodal perception, and task-level agents on an untethered Jetson-powered mobile robot. We use it to study learned/classical skill composition, agent–executor feedback, and resource-aware autonomy in real environments.
Ongoing
Efficient Embodied AI for Intelligent Vehicles and Robots Industry
VLA policies, agent–executor systems, and hardware-aware deployment for real-world vehicle and robot platforms through industry collaboration with Li Auto.
Scaling LawsEfficient VLAAdaptiveWAMLong-horizon VLAcoming soon
⚡ Efficient Foundation Models & Agents
PLM Preprint & Hardware Co-Design Scaling Laws NeurIPS 2026
A 1.8B peripheral language model co-designed with edge hardware, Pareto-optimal scaling laws that choose architectures under latency, memory, and energy constraints (Dancing in Fetters), and RooflineBench for benchmarking on-device LLMs.
PLMCodeScaling LawsRooflineBench🤗 13k+ downloadsDAI 2026
ContextPilot MLSys 2026 & MemoryCraft Under review
Long-context agents, from inference to memory: ContextPilot cuts prefill cost through context reuse, ordering, and deduplication; MemoryCraft is a controlled platform for evaluating agent memory systems jointly with their retrieval regimes, backbones, and token cost.
ContextPilotMemoryCraft
GTA: Efficient Attention Preprint
Grouped-head latent attention: sharing attention maps across head groups and decoding values from a compact latent to cut attention FLOPs and KV-cache size at matched quality.
arXiv
💬 NLP, Agents & AI for Science
GeoGalactica / K2 / GAKG WSDM · AI4X · CIKM
First-generation LLM foundation models for science, covering data acquisition, pre-training, SFT, and RL.
GeoGalacticaK2GAKG 314 stars🤗 2.9k downloads
DS-Agent ICML 2024
Automated data science with LLM agents: case-based reasoning as agent skills and memory in the era of 2024.
arXivCode
Hallucination Detection EMNLP 2023
Uncertainty-based hallucination detection for LLMs with keyword focus and history-aware propagation, improving detection without extra supervision.
arXiv
Platforms. PhysicalAgent mobile base: a self-built wheeled home robot I lead the team on, NVIDIA Jetson Orin on board, lidar, camera, and microphones, ROS 2 navigation, pretrained VLA plus agent loop, running untethered in a real home. Also deployed on Raspberry Pi and consumer phones (PLM), and in-vehicle edge hardware (Li Auto).
-
Luoyang Sun, Guoyang Xia, Fengfa Li, Lei Ren, Xinyu Cui, Haifeng Zhang, Fangxiang Feng, Kaike Zhang, Kun Zhan, Xie Yan, Jun Wang, Cheng Deng*
Conference on Robot Learning, 2026.
-
Under Review Agents · Systems MemoryCraft: How Retrieval Control Reshapes Agent Memory Performance and Cost
Cheng Deng, Eve Sauvage, Danna Zheng, Wenyu Huang, Luo Mai, Mirella Lapata, Jeff Z. Pan
-
Siting Wang, Xiaofeng Wang, Zheng Zhu, Minnan Pei, Xinyu Cui, Cheng Deng, Jian Zhao, Guan Huang, Haifeng Zhang, Jun Wang
-
Luoyang Sun, Jiwen Jiang, Yifeng Ding, Fengfa Li, Yan Song, Haifeng Zhang, Jian Ying, Lei Ren, Kun Zhan, Wei Chen, Yan Xie, Cheng Deng*
Conference on Neural Information Processing Systems, 2026.
-
Yinsicheng Jiang, Yeqi Huang, Liang Cheng, Cheng Deng, Xuan Sun, Luo Mai
International Conference on Machine Learning Systems, 2026.
-
Zhen Bi, Xueshu Chen, Luoyang Sun, Yuhang Yao, Qing Shen, Jungang Lou, Cheng Deng*
International Conference on Distributed Artificial Intelligence, 2026.
-
Siting Wang, Luoyang Sun, Cheng Deng*, Kun Shao, Minnan Pei, Zheng Tian, Haifeng Zhang, Jun Wang
International Conference on Learning Representations, 2026.
-
Luoyang Sun, Cheng Deng, Jiwen Jiang, Xinjian Wu, Haifeng Zhang, Lei Chen, Lionel M. Ni, Jun Wang
-
Cheng Deng*, Luoyang Sun, Jiwen Jiang, Yongcheng Zeng, Xinjian Wu, Wenxin Zhao, Qingfa Xiao, Jiachuan Wang, Haoyang Li, Lei Chen, Lionel M. Ni, Haifeng Zhang, Jun Wang
-
Siyuan Guo, Cheng Deng, Ying Wen, Hechang Chen, Yi Chang, Jun Wang
International Conference on Machine Learning
-
Tianhang Zhang, Lin Qiu, Qipeng Guo, Cheng Deng, Yue Zhang, Zheng Zhang, Chenghu Zhou, Xinbing Wang, Luoyi Fu
Conference on Empirical Methods in Natural Language Processing, 2023.
-
Zhouhan Lin, Cheng Deng, Le Zhou, Tianhang Zhang, Yi Xu, Luoyi Fu, Weinan Zhang, Junxian He, Chao Ma, Yunqiang Zhu, Xinbing Wang, Chenghu Zhou, et al.
-
Cheng Deng, Tianhang Zhang, Zhongmou He, Yi Xu, Qiyuan Chen, Yuanyuan Shi, Luoyi Fu, Weinan Zhang, Xinbing Wang, Chenghu Zhou, Zhouhan Lin, Junxian He
WSDM, 2024
-
Cheng Deng, Yuting Jia, Weinan Zhang, Luoyi Fu, Xinbing Wang, Chenghu Zhou, et al.
CIKM, 2021
* Corresponding author
Services
- Teaching Assistant for Mobile Networks, Program Design (C++), IEEI (B) 2020,2021,2022,2023, and Data Science CS245-2 Kaggle
- Area Chair for ICLR 2027
- Conference Reviewer for ICML, ICLR, NeurIPS, AAAI, CoRL, ACL, EMNLP, AISTATS, ACM MM
- Journal Reviewer for TMLR, TKDE, TMC, FCS, IJGIS, SWJ
- Lead Organizer, Edge AI Agent Workshop, Bayes Centre, University of Edinburgh, 2026
Funding
- Bayes Centre Strategy and Innovation Fellowship, University of Edinburgh, 2025
- Wenjun Wu AI Honour PhD Scholarship, Shanghai Jiao Tong University, 2021
- Presented "Large Discovery Model and Geoscience" at the 3rd International Symposium on Deep Underground Science and Engineering, Glasgow, UK. Aug 2026
- Presented "Efficient AI Agent on the Edge" at the Edge AI Agent Workshop, Bayes Centre, University of Edinburgh (workshop host and organizer). Apr 2026
- Attended the NiklasOPF Podcast, sharing "PLM" and "Hardware co-Design Scaling Law". Feb 2026
- Presented "Efficient LLM on the Edge" at the Cardiff NLP Group seminar in Cardiff University. Feb 2026
- Presented "Efficient Physical AI and Beyond" at the Li Auto AI Sharing seminar. Dec 2025
- Organized and presented a talk on "Advanced Techniques for LLM" and delivered a tutorial on "Full Stack Practice of LLM Training" at RLChina 24, Guangzhou, China. Tutorial materials are available on GitHub. Oct 2024
- Delivered a TED Talk titled "Thinking Outside the Code" at the TEDxNYUShanghai Salon (Theme: Going Meta). Feb 2024
Powered by Jekyll and Minimal Light theme.