Kevin Qu

I am a 1st-year PhD student in the Computer Vision and Learning Group (VLG) at ETH Zürich, advised by Prof. Siyu Tang.

Before I started my PhD, I obtained my Master's in Robotics, Systems and Control at ETH Zürich. During my Master's, I was a Visiting Student Researcher at Stanford University in the Gradient Spaces Lab, advised by Prof. Iro Armeni, working on 3D articulated object modeling. I was also a research intern at the Microsoft Spatial AI Lab under Prof. Marc Pollefeys, where I focused on spatial video understanding with Vision-Language Models (VLMs), and a research assistant at Prof. Konrad Schindler's PRS Lab at ETH, working on diffusion models for dense prediction tasks.

Before that, I obtained a Bachelor's in Electrical and Computer Engineering from the Technical University of Munich. During this time, I was a research intern at the University of Victoria under Prof. Lin Cai, working on communication networks for autonomous driving, and spent a semester abroad at the University of Edinburgh.

Mail  |  GitHub  |  Scholar  |  LinkedIn

headshot
Image Image Image Image Image Image
Research

My current research interests include 3D vision, dynamic scene understanding, and spatial intelligence.

If you share similar interest and would like to collaborate, feel free to reach out!

(* denotes equal contribution)

Image FAMOS: Feed-Forward 3D Articulation Modeling from Sparse Observations
Kevin Qu, Tao Sun, Massimiliano Viola, Liyuan Zhu, Zhizhuo Zhou, Sayan Deb Sarkar, Konrad Schindler, Iro Armeni
arXiv 2026
Paper | Project Page

A feed-forward model that predicts movable-part segmentation and joint parameters from a sparse, unordered set of partial point clouds.

Image Loc3R-VLM: Language-based Localization and 3D Reasoning with Vision-Language Models
Kevin Qu, Haozhe Qi, Mihai Dusmanu, Mahdi Rad, Rui Wang, Marc Pollefeys
arXiv 2026
Paper | Project Page | Code

Equipping 2D Vision-Language Models with 3D spatial understanding capabilities. Inspired by human cognition, we guide the model to learn global scene structure and local viewpoint awareness directly from monocular video.

Image AdaptToken: Entropy-based Adaptive Token Selection for MLLM Long Video Understanding
Haozhe Qi, Kevin Qu, Mahdi Rad, Rui Wang, Alexander Mathis, Marc Pollefeys
arXiv 2026
Paper | Project Page | Code

A training-free framework that turns an MLLM's self-uncertainty into a global control signal for long-video token selection.

Marigold: Affordable Adaptation of Diffusion-Based Image Generators for Image Analysis
Bingxin Ke*, Kevin Qu*, Tianfu Wang*, Nando Metzger*, Shengyu Huang, Bo Li, Anton Obukhov, Konrad Schindler
IEEE Transactions on Pattern Analysis and Machine Intelligence (TPAMI) 2025
Paper | Project Page | Code | Demo

Repurposing text-to-image diffusion models for a range of dense prediction tasks, including monocular depth estimation, surface normal prediction, and intrinsic image decomposition.

Image Marigold-DC: Zero-Shot Monocular Depth Completion with Guided Diffusion
Massimiliano Viola, Kevin Qu, Nando Metzger, Bingxin Ke, Alexander Becker, Konrad Schindler, Anton Obukhov
International Conference on Computer Vision (ICCV) 2025
Paper | Project Page | Code | Demo

Training-free framework for zero-shot depth completion. We use Marigold as an off-the-shelf monocular depth estimator and guide its diffusion process with sparse depth observations.


Last update September 2026. Thanks for the website template .