Ph.D. @ KAUST | Research Scientist Intern @ Qwen Team / Meta AI
I am a Ph.D. candidate in Computer Science at the GenAI Center of Excellence,
KAUST, advised by Prof. Bernard Ghanem. Prior to that, I obtained my Master’s degree from Shanghai Jiao Tong University, and my bachelor’s degree from Xi’an Jiaotong University.
Currently, I am a Research Scientist Intern at
Qwen Team (since July 2026), working on long-horizon agentic multimodal reasoning for Qwen-VL models. Previously, I interned at
Meta AI in 2025 and 2026, focusing on adaptive reasoning and on-policy self-distillation for video-language models.
📢 I am actively seeking full-time positions in 2026! Feel free to reach out: shuming.liu@kaust.edu.sa
Research Interests
I have a broad interest in video understanding and multimodal large language models, with recent work centering on the following:
- Video Reasoning: adaptive auto-thinking video model (VideoAuto-R1)
- Efficient Long-Video Modeling: training-free frame selector (BOLT), parameter-efficient adaptation (AdaTAD)
- Temporal Video Understanding: action localization (CausalTAD), unified detection framework (OpenTAD)
News
2026.07
🔥 I started my internship at Qwen Team as Research Scientist Intern.
2026.04
We released
Neural Computers, a model that simulates a running computer.
2026.04
We released
Tempo, which leverages small VLMs as a temporal compressor for long video understanding.
ECCV 2026 2026.01
We released
VideoAuto-R1, an adaptive auto-thinking model for video understanding.
CVPR 2026 2025.06
I started my internship at Meta AI as Research Scientist Intern.
2025.05
I was awarded the Dean’s List Award of KAUST for 2025 (Top 20%).
2025.02
We released
BOLT, a training-free frame selection method for efficient long-video understanding.
CVPR 2025 2024.07
We released
ColorMAE, a data-independent masking strategy to enhance masked autoencoder pretraining.
ECCV 2024 2024.06
I was awarded the Dean’s List Award of KAUST for 2024 (Top 20%).
2024.05
We released
OpenTAD, a unified framework for temporal action detection that supports 10+ SOTA methods.
CVPRW 2025 2024.02
We released
AdaTAD, a parameter-efficient adaptation method for end-to-end temporal action detection.
CVPR 2024 2024.02
We released
Dr²Net, a dynamic reversible network for finetuning large pretrained models.
CVPR 2024 2024.01
We released
DenoiseLoc, which localizes video activities with precise boundaries via boundary denoising.
ICLR 2024 2023.02
We released
Re²TAL, which rewires pretrained video backbones into reversible networks for end-to-end TAD.
CVPR 2023 2023.02
We released
ETAD, an efficient temporal action detector trained end-to-end with extremely low GPU memory.
CVPRW 2023 Show more news Show less
Selected Publications
-
VideoAuto-R1: Video Auto Reasoning via Thinking Once, Answering Twice
Shuming Liu, Mingchen Zhuge, Changsheng Zhao, Jun Chen, Lemeng Wu, Zechun Liu, and 17 more authors
CVPR 2026
-
BOLT: Boost Large Vision-Language Model Without Training for Long-form Video Understanding
Shuming Liu, Chen Zhao, Tianqi Xu, and Bernard Ghanem
CVPR 2025
-
OpenTAD: A Unified Framework and Comprehensive Study of Temporal Action Detection
Shuming Liu, Chen Zhao, Fatimah Zohra, Mattia Soldan, Alejandro Pardo, Mengmeng Xu, and 7 more authors
CVPR 2025 Workshop on Pixel-level Video Understanding in the Wild
-
Harnessing Temporal Causality for Advanced Temporal Action Detection
Shuming Liu, Lin Sui, Chen-Lin Zhang, Fangzhou Mu, Chen Zhao, and Bernard Ghanem
CVPR 2025 Challenge Report
-
End-to-End Temporal Action Detection with 1B Parameters Across 1000 Frames
Shuming Liu, Chen-Lin Zhang, Chen Zhao, and Bernard Ghanem
CVPR 2024
-
ETAD: Training Action Detection End to End on a Laptop
Shuming Liu, Mengmeng Xu, Chen Zhao, Xu Zhao, and Bernard Ghanem
CVPR 2023 Workshop on Efficient Deep Learning for Computer Vision
-
ICLRW 2026 Neural Computers Mingchen Zhuge, Changsheng Zhao, Haozhe Liu, Zijian Zhou, Shuming Liu, Wenyi Wang, and 13 more authors -
-
-
-
-
-
-
-
A full list of publications is available on Google Scholar.
Research Experience
Ongoing research topic: Long-Horizon Agentic Multimodal Reasoning
China Research topic: On-Policy Self-Distillation for Video-Language Models Dubai, UAE
- Developed an on-policy self-distillation framework that transfers keyframe-aware reasoning from the privileged teacher model to the student video-language model.
- Improved the student's intrinsic keyframe perception without relying on an external frame selector.
Research topic: Adaptive Reasoning for Video-Language Models California, USA
- Developed VideoAuto-R1, one of the first auto-thinking video-language models, reducing inference latency by 70% while improving reasoning accuracy.
- Proposed a "think once, answer twice" training paradigm, coupled with confidence-based early-exit inference.
- Built scalable RL post-training and evaluation infrastructure.
- VideoAuto-R1 was published at CVPR 2026.
Education
2021.09 - 2026.12
Ph.D.,
King Abdullah University of Science and Technology (KAUST), Saudi Arabia.
2018.09 - 2021.04
M.S.,
Shanghai Jiao Tong University (SJTU), China.
2014.09 - 2018.06
B.S.,
Xi'an Jiaotong University (XJTU), China.
Honors and Awards
2025.05
Dean's List Award of KAUST (Top 20%)
2024.06
Dean's List Award of KAUST (Top 20%)
2021.03
Outstanding Graduate of SJTU
2019.12
Scholarship of SJTU (Top 5%)
2018.06
Outstanding Undergraduate of XJTU
2017.12
Scholarship of XJTU (Top 5%)
Service
Conference Reviewer: CVPR, ICCV, ECCV, ICLR, ICML, NeurIPS, AAAI, WACV, BMVC
Journal Reviewer: TPAMI, IJCV, TIP, TMM, Neurocomputing
Teaching Assistant: Introduction to Computer Vision (KAUST), Computer Vision (SJTU)