Zimo He | 何子默

Welcome to my page!

I'm a Ph.D. student at School of Computer Science, Peking University, advised by Prof. Yixin Zhu and Prof. Yizhou Wang. I'm also a research intern at General Vision Lab, BIGAI, working on general 3D vision and humanoid robot learning. I earned my Bachelor's degree from Yuanpei College, Peking University, majoring in Artificial Intelligence.

Email  /  GitHub

profile photo

Research

My research interests lie in the intersection of 3D vision, robotics and computer graphics, including human-scene interaction, humanoid robot and human motion synthesis. My long-term goal is to build a general embodied AI that can see, understand and interact with environments, both in virtual and real world.

Image UniAct: Unified Motion Generation and Action Streaming for Humanoid Robots
Nan Jiang*, Zimo He*, Wanhe Yu, Lexi Pang, Yunhao Li, Hongjie Li, Jieming Cui, Yuhan Li, Yizhou Wang, Yixin Zhu†, Siyuan Huang
Under Review
project page / arXiv / code / video

We introduce a unified framework that enables humanoid robots to execute multimodal instructions by generating motion tokens through a shared representation and achieving real-time whole-body control.

Image MotionMaster: Generalizable Text-Driven Motion Generation and Editing
Nan Jiang*, Yunhao Li*, Lexi Pang*, Zimo He, Siyuan Huang†, Yixin Zhu
CVPR, 2026
project page / paper / code

By finetuning a pretrained multimodal LLM on large-scale motion data, MotionMaster unifies text-driven motion generation and editing in an end-to-end framework, achieving strong zero-shot generalization across multi-action composition and fine-grained body-part editing.

Image Dynamic Motion Blending for Versatile Motion Editing
Nan Jiang*, Hongjie Li*, Ziye Yuan*, Zimo He, Yixin Chen, Tengyu Liu, Yixin Zhu†, Siyuan Huang
CVPR, 2025
project page / arXiv / code

Our text-guided motion editor combines dynamic body-part blending augmentation with an auto-regressive diffusion model to enable diverse motion editing without extensive pre-collected training data.

Image
Autonomous Character-Scene Interaction Synthesis from Text Instruction
Nan Jiang*, Zimo He*, Zi Wang, Hongjie Li, Yixin Chen, Siyuan Huang†, Yixin Zhu
SIGGRAPH Asia, 2024
project page / arXiv / code

We propose a framework synthesizing multi-stage scene-aware human motions autonomously and directly from the text instruction and goal location. We also introduce LINGO, a comprehensive language-annotated MoCap dataset, featuring Human-Scene Interaction (HSI) motions.

Teaching

Cognitive Reasoning (TA) Fall 2026
Cognitive Reasoning (TA) Fall 2025
Directed Research in AI System (I) (TA) Spring 2025
Directed Research in AI System (I) (TA) Fall 2024

Experience

Image School of Computer Science, Peking University
Aug. 2025 - Jul. 2030 (expected)

Ph.D. Student
Advisor: Prof. Yixin Zhu and Prof. Yizhou Wang
Image Beijing Institute for General Artificial Intelligence (BIGAI)
Apr. 2024 - Present

Research Intern
Advisor: Dr. Siyuan Huang and Dr. Tengyu Liu
Image Cognitive Reasoning (CoRe) Lab, Institute for Artificial Intelligence, Peking University
Jul. 2023 - Present

Research Assistant
Advisor: Prof. Yixin Zhu
Image Tong Class, Yuanpei College, Peking University
Aug. 2021 - Jul. 2025

Undergraduate Student

This website is a modification of Jon Barron's website.