I was a Research Scientist and technical lead for cross-functional projects at Stability AI, working on structured generative models and multimodal systems for image, video, 3D, and dynamic world modeling.
My research focuses on structured and controllable generative models for interactive multimodal generation across image, video, and 3D. I study how to learn structured representations of dynamic scenes, model temporal dynamics, and build controllable generative systems with multimodal inputs.
My work spans 3D perception/reconstruction/generation, controllable video generation, and multimodal generative systems, with an emphasis on temporal consistency, interaction, and structured scene understanding. In the long term, I am interested in building persistent and interactive generative systems that connect multimodal generation, dynamics modeling, and world understanding.
3D Perception, Reconstruction, and Dynamic Generation:
SurMo,
EgoRenderer,
HVTR,
HVTR++
Learning structured 3D representations from visual observations and modeling temporal dynamics for reconstruction, animation, and dynamic scene generation.
Structured Dynamic Scene Modeling and Control:
StructLDM,
FashionEngine,
HumanLiff
Generative models using structured state representations and multimodal inputs for controllable and interactive generation.
Motion Modeling and Controllable Video Dynamics:
HumANDiff,
SHV4D
Video generation with motion consistency and intrinsic control over pose and camera.
Human-Centric World Modeling for Stable Human-Object-Scene Interaction Synthesis.
Tao Hu et al. (Research Lead & Technical Lead)
→An interactive world model using a hybrid framework combining 3D control and video diffusion for controllable human-object-scene interaction synthesis. Developed a human-centric world model from structured states, geometry, and multimodal conditioning for stable video generation.
Tiantian Wang, Chun-Han Yao, Tao Hu, Mallikarjun Byrasandra Ramalinga Reddy, Ming-Hsuan Yang, Varun Jampani.
The 37th British Machine Vision Conference (BMVC), 2026 [Paper] →A generalizable framework that leverages generic video foundation models for controllable 4D human generation from a single image, enabling pose and camera control.
Tao Hu, Fangzhou Hong, Zhaoxi Chen, Ziwei Liu.
arXiv:2404.01655, Technical Report [Project Page] [Video] [arXiv] → A unified framework for interactive 3D human generation and editing that leverages structured multimodal pretraining to support diverse controls from text, images, and hand-drawn sketches.
Tao Hu, Fangzhou Hong, Ziwei Liu.
European Conference on Computer Vision (ECCV 2024)
[Project Page] [Video] [Code] [arXiv] [Media Coverage]
[Media Coverage in Chinese: 1,2]
→A new paradigm for 3D human generation via structured latent representations learned through pretraining of a human prior for controllable diffusion modeling.
Shoukang Hu, Fangzhou Hong, Tao Hu , Liang Pan, Weiye Xiao, Haiyi Mei, Lei Yang, Ziwei Liu
International Journal of Computer Vision (IJCV 2025) [Paper][Project Page][Code] → A diffusion-based approach for layer-wise controllable 3D human generation.
Tao Hu, Fangzhou Hong, Ziwei Liu.
IEEE Conference on Computer Vision and Pattern Recognition (CVPR 2024)
[Paper]
[Project Page]
[Video]
[Code]
[Media Coverage in Chinese: Media Heart, SenseTime Research]
→ A novel framework for 4D human motion modeling that jointly learns temporal motion dynamics and human appearance representations from videos through a surface-based triplane representation.
Tao Hu, Hongyi Xu, Linjie Luo, Tao Yu, Zerong Zheng, He Zhang, Yebin Liu, Matthias Zwicker.
IEEE Transactions on Visualization and Computer Graphics (TVCG 2023)
[Paper]
[Project Page]
[Video]
[Code]
→A virtual teleportation system using sparse view cameras based on a novel structured texel-aligned multimodal representation.
Tao Hu, Tao Yu, Zerong Zheng, He Zhang, Yebin Liu, Matthias Zwicker.
International Conference on 3D Vision (3DV 2022)
[Paper]
[Project Page] [Video] [Poster] [arXiv]
[Code] → A novel hybrid neural rendering framework that combines classical volumetric rendering and probabilistic generative models for efficient and realistic human avatar rendering.
Tao Hu, Geng Lin, Zhizhong Han, Matthias Zwicker.
IEEE Winter Conference on Applications of Computer Vision (WACV 2021) [Paper] [Code] [arXiv] → Extend the multi-view representation for generalizable geometry/texture reconstructions from single RGB images.
Tao Hu, Zhizhong Han, Matthias Zwicker.
AAAI Conference on Artificial Intelligence (AAAI 2020, Oral, top 10% among accepted papers in 3D vision track)
[Paper] [Code] [arXiv] → Introduce a self-supervised multi-view consistent inference technique to enforce geometric consistency for multi-view representation.
Tao Hu, Zhizhong Han, Abhinav Shrivastava, Matthias Zwicker.
IEEE ICCV Geometry Meets Deep Learning Workshop (ICCVW 2019, Oral) [Paper] [Code] [arXiv] → Present multi-view based 3D shape representation with a multi-view completion net for dense 3D shape completion.
Tao Hu, Gangyi Ding, Lijie Li, Longfei Zhang.
Highlights of Sciencepaper, Chinese Journal, May 2016. →Propose a parallel video player plugin for CryEngine3 for a speedup from 16 FPS to 54 FPS at a large-scale virtual stage with 40 LED screens playing videos simultaneously for digital performance.
Reviewer for CVPR, ICCV, ECCV, NeurIPS, ICML, ICLR, AAAI, CFG, CVIU, 3DV, VR, etc.
Selected Awards & Honors
Graduate National Scholarship (Top 2%), Ministry of Education of China 2016
Undergraduate National Scholarship (Top 2%), Ministry of Education of China 2014
Teaching Experience
Teaching Assistant, Dept. of Computer Science, UMD.
CMSC425 Game Programming (Prof. Roger Eastman), Fall 2019
CMSC425 Game Programming (Prof. Roger Eastman), Spring 2019
CMSC 216 Introduction to Computer Systems (Mr. Laurence Herman), Fall 2018