Shuming Liu

prof_pic.jpg

King Abdullah University of Science and Technology (KAUST)

Ph.D. @ KAUST | Research Scientist Intern @ Qwen Team / Meta AI

I am a Ph.D. candidate in Computer Science at the GenAI Center of Excellence, KAUST, advised by Prof. Bernard Ghanem. Prior to that, I obtained my Master’s degree from Shanghai Jiao Tong University, and my bachelor’s degree from Xi’an Jiaotong University.

Currently, I am a Research Scientist Intern at Qwen Team (since July 2026), working on long-horizon agentic multimodal reasoning for Qwen-VL models. Previously, I interned at Meta AI in 2025 and 2026, focusing on adaptive reasoning and on-policy self-distillation for video-language models.

📢 I am actively seeking full-time positions in 2026! Feel free to reach out: shuming.liu@kaust.edu.sa

Research Interests

I have a broad interest in video understanding and multimodal large language models, with recent work centering on the following:

  • Video Reasoning: adaptive auto-thinking video model (VideoAuto-R1)
  • Efficient Long-Video Modeling: training-free frame selector (BOLT), parameter-efficient adaptation (AdaTAD)
  • Temporal Video Understanding: action localization (CausalTAD), unified detection framework (OpenTAD)

News

2026.07
🔥 I started my internship at Qwen Team as Research Scientist Intern.
2026.04
We released Neural Computers, a model that simulates a running computer.
2026.04
We released Tempo, which leverages small VLMs as a temporal compressor for long video understanding. ECCV 2026
2026.01
We released VideoAuto-R1, an adaptive auto-thinking model for video understanding. CVPR 2026
2025.06
I started my internship at Meta AI as Research Scientist Intern.
2025.05
I was awarded the Dean’s List Award of KAUST for 2025 (Top 20%).
2025.02
We released BOLT, a training-free frame selection method for efficient long-video understanding. CVPR 2025
2024.07
We released ColorMAE, a data-independent masking strategy to enhance masked autoencoder pretraining. ECCV 2024
2024.06
We won 4 championships in CVPR 2024 Challenges, including Action Recognition, Action Detection, Audio-Based Interaction Detection, and Moment Queries!
2024.06
I was awarded the Dean’s List Award of KAUST for 2024 (Top 20%).
2024.05
We released OpenTAD, a unified framework for temporal action detection that supports 10+ SOTA methods. CVPRW 2025
2024.02
We released AdaTAD, a parameter-efficient adaptation method for end-to-end temporal action detection. CVPR 2024
2024.02
We released Dr²Net, a dynamic reversible network for finetuning large pretrained models. CVPR 2024
2024.01
We released DenoiseLoc, which localizes video activities with precise boundaries via boundary denoising. ICLR 2024
2023.02
We released Re²TAL, which rewires pretrained video backbones into reversible networks for end-to-end TAD. CVPR 2023
2023.02
We released ETAD, an efficient temporal action detector trained end-to-end with extremely low GPU memory. CVPRW 2023
Show more news Show less

Selected Publications

  1. videoauto_r1.png
    VideoAuto-R1: Video Auto Reasoning via Thinking Once, Answering Twice
    Shuming Liu, Mingchen Zhuge, Changsheng Zhao, Jun Chen, Lemeng Wu, Zechun Liu, and 17 more authors
    CVPR 2026
  2. bolt.png
    BOLT: Boost Large Vision-Language Model Without Training for Long-form Video Understanding
    Shuming Liu, Chen Zhao, Tianqi Xu, and Bernard Ghanem
    CVPR 2025
  3. opentad.png
    OpenTAD: A Unified Framework and Comprehensive Study of Temporal Action Detection
    Shuming Liu, Chen Zhao, Fatimah Zohra, Mattia Soldan, Alejandro Pardo, Mengmeng Xu, and 7 more authors
    CVPR 2025 Workshop on Pixel-level Video Understanding in the Wild
  4. causaltad.png
    Harnessing Temporal Causality for Advanced Temporal Action Detection
    Shuming Liu, Lin Sui, Chen-Lin Zhang, Fangzhou Mu, Chen Zhao, and Bernard Ghanem
    CVPR 2025 Challenge Report
  5. adatad.png
    End-to-End Temporal Action Detection with 1B Parameters Across 1000 Frames
    Shuming Liu, Chen-Lin Zhang, Chen Zhao, and Bernard Ghanem
    CVPR 2024
  6. etad.png
    ETAD: Training Action Detection End to End on a Laptop
    Shuming Liu, Mengmeng Xu, Chen Zhao, Xu Zhao, and Bernard Ghanem
    CVPR 2023 Workshop on Efficient Deep Learning for Computer Vision
  1. ICLRW 2026 Neural Computers
    Mingchen Zhuge, Changsheng Zhao, Haozhe Liu, Zijian Zhou, Shuming Liu, Wenyi Wang, and 13 more authors
  2. ECCV 2026 Small Vision-Language Models are Smart Compressors for Long Video Understanding
    Junjie Fei, Jun Chen, Zechun Liu, Yunyang Xiong, Chong Zhou, Wei Wen, and 10 more authors
  3. CVPR 2026 Mixture of States: Routing Token-Level Dynamics for Multimodal Generation
    Haozhe Liu, Ding Liu, Mingchen Zhuge, Zijian Zhou, Tian Xie, Sen He, and 12 more authors
  4. arXiv 2025 TimeLoc: A Unified End-to-End Framework for Precise Timestamp Localization in Long Videos
    Chen-Lin Zhang, Lin Sui, Shuming Liu, Fangzhou Mu, Zhangcheng Wang, and Bernard Ghanem
  5. ECCV 2024 ColorMAE: Exploring data-independent masking strategies in Masked AutoEncoders
    Carlos Hinojosa, Shuming Liu, and Bernard Ghanem [Code]
  6. CVPR 2024 Dr²Net: Dynamic Reversible Dual-Residual Networks for Memory-Efficient Finetuning
    Chen Zhao, Shuming Liu, Karttikeya Mangalam, Guocheng Qian, Fatimah Zohra, Abdulmohsen Alghannam, and 2 more authors
  7. CVPRW 2024 Look, Listen, and Attack: Backdoor Attacks Against Video Action Recognition
    Hasan Hammoud, Shuming Liu, Mohammed Alkhrashi, Fahad AlBalawi, and Bernard Ghanem
  8. ICLR 2024 Boundary-Denoising for Video Activity Localization
    Mengmeng Xu, Mattia Soldan, Jialin Gao, Shuming Liu, Juan-Manuel Perez-Rua, and Bernard Ghanem [Code]
  9. CVPR 2023 Re²TAL: Rewiring Pretrained Video Backbones for Reversible Temporal Action Localization
    Chen Zhao, Shuming Liu, Karttikeya Mangalam, and Bernard Ghanem [Code]

A full list of publications is available on Google Scholar.

Research Experience

Qwen Team - Research Scientist Intern 2026.07 - Present

Ongoing research topic: Long-Horizon Agentic Multimodal Reasoning

China
Meta AI - Research Scientist Intern 2026.04 - 2026.06
Research topic: On-Policy Self-Distillation for Video-Language Models Dubai, UAE
  • Developed an on-policy self-distillation framework that transfers keyframe-aware reasoning from the privileged teacher model to the student video-language model.
  • Improved the student's intrinsic keyframe perception without relying on an external frame selector.
Meta AI - Research Scientist Intern 2025.05 - 2025.10
Research topic: Adaptive Reasoning for Video-Language Models California, USA
  • Developed VideoAuto-R1, one of the first auto-thinking video-language models, reducing inference latency by 70% while improving reasoning accuracy.
  • Proposed a "think once, answer twice" training paradigm, coupled with confidence-based early-exit inference.
  • Built scalable RL post-training and evaluation infrastructure.
  • VideoAuto-R1 was published at CVPR 2026.

Education

2021.09 - 2026.12
Ph.D., King Abdullah University of Science and Technology (KAUST), Saudi Arabia.
2018.09 - 2021.04
M.S., Shanghai Jiao Tong University (SJTU), China.
2014.09 - 2018.06
B.S., Xi'an Jiaotong University (XJTU), China.

Honors and Awards

2025.05
Dean's List Award of KAUST (Top 20%)
2024.06
Dean's List Award of KAUST (Top 20%)
2021.03
Outstanding Graduate of SJTU
2019.12
Scholarship of SJTU (Top 5%)
2018.06
Outstanding Undergraduate of XJTU
2017.12
Scholarship of XJTU (Top 5%)

Service

Conference Reviewer: CVPR, ICCV, ECCV, ICLR, ICML, NeurIPS, AAAI, WACV, BMVC

Journal Reviewer: TPAMI, IJCV, TIP, TMM, Neurocomputing

Teaching Assistant: Introduction to Computer Vision (KAUST), Computer Vision (SJTU)