Bin Xiao
Member of Technical Staff
Microsoft Superintelligence Team · Redmond, WA
I am an AI researcher studying multimodal foundation models, computer vision, and agentic AI. My work spans vision-language model training and post-training, visual recognition, document understanding, and human pose estimation.
Research interests
- Multimodal foundation models. Unified vision-language representations for captioning, OCR, grounding, and document understanding.
- Reasoning and agentic systems. Post-training for reasoning, coding, tool use, and multi-step task completion.
- Efficient models. Compact multimodal models for low-latency and resource-constrained settings.
- Visual representation learning. High-resolution networks, convolutional vision transformers, and human pose estimation.
Selected publications & reports
A selection of my research. Full publication list on Google Scholar.
Peer-reviewed publications
-
Florence-2: Advancing a Unified Representation for a Variety of Vision Tasks
A prompt-based sequence-to-sequence model for captioning, object detection, grounding, and segmentation.
BibTeX
@InProceedings{Xiao_2024_CVPR, author = {Xiao, Bin and Wu, Haiping and Xu, Weijian and Dai, Xiyang and Hu, Houdong and Lu, Yumao and Zeng, Michael and Liu, Ce and Yuan, Lu}, title = {Florence-2: Advancing a Unified Representation for a Variety of Vision Tasks}, booktitle = {Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)}, year = {2024}, pages = {4818-4829} } -
CvT: Introducing Convolutions to Vision Transformers
Convolutional token embeddings and projections introduce spatial inductive biases into a hierarchical vision transformer.
BibTeX
@InProceedings{Wu_2021_ICCV, author = {Wu, Haiping and Xiao, Bin and Codella, Noel and Liu, Mengchen and Dai, Xiyang and Yuan, Lu and Zhang, Lei}, title = {CvT: Introducing Convolutions to Vision Transformers}, booktitle = {Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV)}, year = {2021}, pages = {22-31} } -
Deep High-Resolution Representation Learning for Human Pose Estimation
HRNet maintains high-resolution representations throughout the network through parallel multi-resolution branches and repeated feature fusion.
BibTeX
@InProceedings{Sun_2019_CVPR, author = {Sun, Ke and Xiao, Bin and Liu, Dong and Wang, Jingdong}, title = {Deep High-Resolution Representation Learning for Human Pose Estimation}, booktitle = {Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)}, year = {2019}, pages = {5693-5703} } -
Simple Baselines for Human Pose Estimation and Tracking
Simple baseline methods for human pose estimation and tracking, using a deconvolutional head for keypoint prediction.
Technical reports & model releases
-
MAI-Thinking-1 and MAI-Code-1-Flash
Reasoning and coding models for mathematical reasoning, software engineering, and tool-based coding workflows.
-
Phi-3 Vision / Phi-3.5 Vision
Compact vision-language models for image understanding, visual reasoning, OCR, and document understanding.
Research background
-
2025 – present
Post-training for reasoning, coding, and agentic models, including planning and tool use.
-
2024 – 2025
Multimodal Llama post-training for visual understanding.
-
2024
Led Phi-3 Vision and Phi-3.5 Vision, compact multimodal language models.
-
2020 – 2023
Led the Florence-1 and Florence-2 vision foundation models.
-
2018 – 2021
Developed visual representations and pose estimation methods in CvT, HRNet, and Simple Baselines.
Recognition
Selected challenge results.
-
2019
1st place, Look into Person Challenge, Single-Person Pose Estimation.
-
2019
2nd place, Object365 Challenge, Full track.
-
2018
1st place, PoseTrack Multi-Person Pose Tracking Challenge.
-
2018
2nd place, COCO Keypoint Detection Challenge.
Citation-based recognition: HRNet (CVPR 2019), CvT (ICCV 2021), and Simple Baselines (ECCV 2018) appear in Paper Digest's most-influential-paper lists (March 2026).
Contact & profiles
For professional correspondence, please connect via LinkedIn.
- Scholar
- Google Scholar
- ORCID
- 0000-0001-6477-5911
- GitHub
- github.com/leoxiaobin
- Hugging Face
- huggingface.co/leoxiaobin
- linkedin.com/in/leobinxiao
- Location
- Microsoft Superintelligence Team · Redmond, WA