About Me

I am a PhD candidate (2024–) in the Gradient Spaces Group at Stanford University, advised by Prof. Iro Armeni, and part of the Stanford Vision Lab. During my PhD, I have interned at Waymo Perception (Summer ’26), working on streaming temporal reasoning, and at Microsoft Spatial AI Lab (Summer ’25), working on efficient video tokenization. I am partly supported by a Google XR grant on 4D understanding.

My research focuses on spatial intelligence and multimodal video understanding. I build efficient visual representations that capture space and time, align real-world 3D scenes across modalities, and enable controllable built-environment generation.

Before PhD, I received my M.Sc. in Computer Science from ETH Zurich, advised by Prof. Marc Pollefeys, during which I interned at Qualcomm XR working on real-time SLAM. Earlier, I was a computer vision research engineer at Mercedes-Benz R&D and spent a year in Prof. Vincent Lepetit's lab at TU Graz working on hand-object pose estimation.

Interests

  • Spatial Intelligence
  • Multimodal Video Understanding
  • Visual Representation Learning

News

I joined Waymo Perception (Mountain View) as a research intern, working on streaming temporal reasoning.

Recognized as an Outstanding Reviewer at CVPR 2026. Very proud of this!

Invited talks on “Codec Primitives for Efficient Video Understanding” at Google DeepMind and Valeo AI in Paris.

CoPE-VideoLM is out: efficient codec-aware tokenization for video understanding.

Invited talk on GuideFlow3D at Voxel51 Best of NeurIPS.

GuideFlow3D was accepted to NeurIPS 2025. See you in San Diego!

Invited talks on “Scalable Cross-Modal 3D Scene Understanding” at Google XR Research and Imagine Lab.

I joined Microsoft Spatial AI Lab (Zurich) as a research intern, working on efficient video tokenization.

Show moreShow less

CrossOver was accepted to CVPR 2025 as a Highlight. See you in Nashville!

Career update I joined Stanford University as a PhD student in Computer Vision.

I joined Qualcomm XR (Amsterdam) as a research intern, working on real-time SLAM for extended reality.

SGAligner was accepted to ICCV 2023. See you in Paris!

I started my MSc in Computer Science at ETH Zurich.

Keypoint Transformer was accepted to CVPR 2022 as an Oral. See you in New Orleans!

I joined Mercedes-Benz R&D as a Computer Vision Research Engineer.

Monte Carlo Scene Search was accepted to CVPR 2021.

I started as a Research Assistant at IVC, TU Graz with Prof. Vincent Lepetit, supported by a Qualcomm fellowship.

Invited Talks

Research

Visual Representation Learning

Efficient visual encoders for spatial and video understanding.

CoPE-VideoLM

Controllable Generation

Guiding pre-trained generative models to control shape and appearance.

SpaceFlowGuideFlow3D

3D Scene Understanding

Understanding real-world environments, from alignment to affordances.

CrossOverSGAlignerFAMOS

Publications

Google Scholar

* Equal contribution · † Equal supervision

2026

PreprintUnder review

SpaceFlow: Locally Controllable 3D Generation

Neil De La Fuente*, Joan Lafuente Baeza*, Mukhammadali Sayfiddinov*, Felicia Scharitzer*, Marc Pollefeys, Ata Çelen, Sayan Deb Sarkar†, Elisabetta Fedele†

Training-free 3D generation from geometric primitives, with per-part control over shape adherence and part-specific text or image appearance cues.

2025

2023

2022

2021

MCSS
CVPR 2021

Monte Carlo Scene Search for 3D Scene Understanding

Shreyas Hampali*, Sinisa Stekovic*, Sayan Deb Sarkar, Chetan Srinivasa Kumar, Friedrich Fraundorfer, Vincent Lepetit

Monte-Carlo Tree Search (MCTS) based analysis-by-synthesis method to recover complete scene (3D layout+objects) from a noisy RGB-D scan.

2020

Background

Experience

Education

Supervision

Academic Services

Get in touch

Let’s talk research.

I’m always open to research collaborations. If you’re around the Bay Area, let’s grab a coffee!

sdsarkar@stanford.edu