Updates
- [2026-09] Our surgical video annotation platform LabelSurg, powered by our state-of-the-art surgical AI models, is available at this link.
- [2026-09] Released the GraSP subset of SurgΣ-DB on Hugging Face; with new task annotations including triplets, safety assessment, and more, along with the corresponding annotation protocol.
- [2026-04] Released the CholecT50 subset of SurgΣ-DB on Hugging Face.
- [2026-03] Our paper SurgΣ: A Spectrum of Large-Scale Multimodal Data and Foundation Models for Surgical Intelligence is available on arXiv.
Abstract
Surgical intelligence has the potential to improve the safety and consistency of surgical care, yet most existing surgical AI frameworks remain task-specific and struggle to generalize across procedures and institutions. Although multimodal foundation models, particularly multimodal large language models, have demonstrated strong cross-task capabilities across various medical domains, their advancement in surgery remains constrained by the lack of large-scale, systematically curated multimodal data. To address this challenge, we introduce SurgΣ, a spectrum of large-scale multimodal data and foundation models for surgical intelligence. At the core of this framework lies SurgΣ-DB, a large-scale multimodal data foundation designed to support diverse surgical tasks. SurgΣ-DB consolidates heterogeneous surgical data sources (including open-source datasets, curated in-house clinical collections and web-source data) into a unified schema, aiming to improve label consistency and data standardization across heterogeneous datasets. SurgΣ-DB spans 6 clinical specialties and diverse surgical types, providing rich image- and video-level annotations across 18 practical surgical tasks covering understanding, reasoning, planning, and generation, at an unprecedented scale (over 5.98M conversations). Beyond conventional multimodal conversations, SurgΣ-DB incorporates hierarchical reasoning annotations, providing richer semantic cues to support deeper contextual understanding in complex surgical scenarios. We further provide empirical evidence through recently developed surgical foundation models built upon SurgΣ-DB, illustrating the practical benefits of large-scale multimodal annotations, unified semantic design, and structured reasoning annotations for improving cross-task generalization and interpretability.
Examples of Tasks in SurgΣ-DB
Overview of the Data Curation and Annotation Pipeline
Summary of Currently Released Data
For the ground-truths we labeled, we also provide corresponding annotation guidelines (please click ✓ for more details).
| Datasets | Understanding & Reasoning | Planning & Generation | ||||||||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| 1 | 2 | 3 | 4 | 5 | 6 | 7 | 8 | 9 | 10 | 11 | 12 | 13 | 14 | 15 | 16 | 17 | 18 | |
| CholecT50 | ✓ | ✓ | ✓ | ✓ | ✓ | NA | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | |
| GraSP | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ||
Understanding and Reasoning: 1: Instrument Recognition; 2: Instrument Localization; 3: Instrument Segmentation; 4: Tissue and Organ Recognition; 5: Tissue and Organ Localization; 6: Phase Recognition; 7: Step Recognition; 8: Action Recognition; 9: Triplet Recognition; 10: Depth Estimation; 11: Safety Assessment; 12: Surgical Image Captioning; 13: Surgical Video Captioning.
Planning and Generation: 14: Action Remaining Prediction; 15: Next Action Planning; 16: Desmoking; 17: Next Frame Prediction; 18: Conditional Surgical Video Generation.
BibTeX
@article{zeng2026surgsigma,
title={Surg{$\Sigma$}: A Spectrum of Large-Scale Multimodal Data and Foundation Models for Surgical Intelligence},
author={Zhitao Zeng and Mengya Xu and Jian Jiang and Pengfei Guo and Yunqiu Xu and Zhu Zhuo and Chang Han Low and Yufan He and Dong Yang and Chenxi Lin and Yiming Gu and Jiaxin Guo and Yutong Ban and Daguang Xu and Qi Dou and Yueming Jin},
journal={arXiv preprint arXiv:2603.16822},
year={2026}
}