I study intelligence in the physical world by structuring raw observations into spatial code. My works laid the foundation, spanning vision encoders and world models. I am best known for TransUNet, which unifies local and global context and has over 10,000 citations.
I am honored as a Young American Scientist by the pretigious magazine Scientific American, a Siebel Scholar in bioengineering, and a recipient of a few best paper awards.
07/2026. I am deeply honored to be recognized by Scientific American as one of the Young American Scientists. The magazine has inspired my curiosity about science since childhood.
06/2026. Co-organized WMAS (World Model Meets Active Sensing) and GCV (Generative Computer Vision) workshops at IEEE/CVF Computer Vision and Pattern Recognition Conference (CVPR) 2026.
06/2026. Invited to serve as Area Chair for the 41st AAAI Conference on Artificial Intelligence (AAAI-27).
04/2026. World-in-World is selected as Oral presentation (1%) in ICLR 2026.
Happy to chat if you’re looking for research collaboration or research opportunities.
Research
When we watch a ball bounce, we do more than recognize pixels. We infer 3D structure, anticipate what may happen next, and reason about physical interactions over time. What representation could give machines the same ability to understand and act in a dynamic world?
I seek fundamental representations of the physical-world observations. My work builds increasingly rich world priors through learned multimodal representations (Chen ’21, ’24, ’26) and generative representations (Lu ’25; Chen ’26). This research path aims to enable machines to reason, plan and adapt through closed-loop interaction (Zhang ’26).
Beyond representations, I apply these ideas to robots operating in complex environments and to cancer diagnosis and treatment, with the long-term goal of improving the lives of patients and their families.
Scientific American, July/August 2026Illustration for the feature
Shanshan Zhong, SYSU MS → CMU LTI PhD 1 publication on 4D-Animal.
Acknowledgement
My doctoral research was made possible through the generous support of ARL, IARPA, NSF, NIH, ONR, Lambda, NVIDIA, Google Cloud, JHU, Stanford, Harvard, the Biswas Foundation, the Siebel Foundation, the Patrick J. McGovern Foundation, and the Lustgarten Foundation.