Ph.D. Student in Computer Science
Yale University
I am looking for a full-time position starting May 2027 (earlier also possible)!
I'm a Ph.D. student in Computer Science (2023 - [Expected] 2027) at Yale University, advised by Prof. Yuval Kluger (primary) and Prof. Daniel Rakita. Previous to that, I obtained my B.Eng. in Computer Science (2019 - 2023) at ShanghaiTech University, minor in Innovation and Entrepreneurship. During my PhD, I did research internship at Nvidia Research and Meta Reality Lab.
I conduct research on Multimodal Learning for Agents and Robotics, especially Vision-Language Models (VLMs), LLM Agents, Diffusion Models, World Models, and 3D Vision. My line of work in "Language for 3D Vision" explores how vision-language models can perceive and understand the world like humans do.
I was also fortunate to intern at Shanghai AI Lab and UISEE during my undergraduate studies.
CVPR (2022, 2025 [Outstanding Reviewer], 2026), ICCV (2023, 2025), ECCV (2024, 2026), ICML (2025, 2026), ICLR (2025, 2026), NeurIPS (2024, 2025), BMVC (2026), ACM MM (2023, 2025), AISTATS (2024, 2025), ICASSP (2024, 2025), TCSVT (journal), TMLR (journal)
Multimodal Representation Learning towards Agentic AI
Since the age of 13, deeply touched by Foundation by Isaac Asimov, my dream has been to create an AI who can think like humans (just like the dream of Prof. Jürgen Schmidhuber). When humans perceive the surrounding environment, we see (vision), hear (audio), feel (tactile) simultaneously to reason (language) and understand (neural signal), then interact (action) with the world. Therefore, I conduct research on Multimodal Representation Learning for Agentic AI, which spans both embodied agents (robotics) and digital agents (LLM/VLM-based agents). Despite differing in their input modalities and action spaces, they are unified by a common goal: the need to perceive, reason, and act intelligently within their environments.
My key reserach questions are (1) how to obtain good multimodal representations for agents and robots, and (2) how to use those representations to make agents and robots better perceive, reason, understand, and interact with the world. Specifically, my line of work in "Language for 3D Vision" (DepthCLIP, PointCLIPv2, WorDepth, RSA, Iris) explores how vision-language models can perceive and understand the world like humans do.
First/Co-first author publications are highlighted.
(* indicates equal contributions)
Under Review of Nature | project page, code, paper
arXiv technical report, 2026 | NE Agents Day 2026 | project page, code, paper
ECCV 2026 (Oral) | project page, code, paper
arXiv technical report, 2026 | project page, code, paper
arXiv technical report, 2026 | paper
Under Review
Under review of ACM TOMM
arXiv technical report, 2024
ACM Multimedia 2022, accepted as Brave New Idea (Accepte Rate<=12.5%) | code
arXiv technical report, 2021
Final Project of CS181 Artificial Intelligence, 2021 Fall, ShanghaiTech University | code
Final Project of CS282 Machine Learning, 2021 Spring, ShanghaiTech University
Final Project of CS280 Deep Learning, 2020 Fall, ShanghaiTech University
I'm an amateur Unity game developer, previous supervised by Brain Cox, screenshots of my previous works have been shown below.
I'm also an amateur pianist, trombone player, guitar player, and Chinese folk singer.
I have been playing Tarot since 2014, dedicating to combining Tarot with modern psychology to serve as a tool for consciousness.
Previously, I volunteered at WWF-China and Greenpeace