Bin Xiao

Member of Technical Staff
Microsoft Superintelligence Team · Redmond, WA

I am an AI researcher studying multimodal foundation models, computer vision, and agentic AI. My work spans vision-language model training and post-training, visual recognition, document understanding, and human pose estimation.

Research interests

  • Multimodal foundation models. Unified vision-language representations for captioning, OCR, grounding, and document understanding.
  • Reasoning and agentic systems. Post-training for reasoning, coding, tool use, and multi-step task completion.
  • Efficient models. Compact multimodal models for low-latency and resource-constrained settings.
  • Visual representation learning. High-resolution networks, convolutional vision transformers, and human pose estimation.

Selected publications & reports

A selection of my research. Full publication list on Google Scholar.

  • Florence-2: Advancing a Unified Representation for a Variety of Vision Tasks

    Bin Xiao, Haiping Wu, Weijian Xu, Xiyang Dai, Houdong Hu, Yumao Lu, Michael Zeng, Ce Liu, Lu Yuan

    CVPR 2024Oral presentation

    A prompt-based sequence-to-sequence model for captioning, object detection, grounding, and segmentation.

    BibTeX
    @InProceedings{Xiao_2024_CVPR,
      author = {Xiao, Bin and Wu, Haiping and Xu, Weijian and Dai, Xiyang and Hu, Houdong and Lu, Yumao and Zeng, Michael and Liu, Ce and Yuan, Lu},
      title = {Florence-2: Advancing a Unified Representation for a Variety of Vision Tasks},
      booktitle = {Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)},
      year = {2024},
      pages = {4818-4829}
    }
  • CvT: Introducing Convolutions to Vision Transformers

    Haiping Wu, Bin Xiao, Noel Codella, Mengchen Liu, Xiyang Dai, Lu Yuan, Lei Zhang

    ICCV 2021

    Convolutional token embeddings and projections introduce spatial inductive biases into a hierarchical vision transformer.

    BibTeX
    @InProceedings{Wu_2021_ICCV,
      author = {Wu, Haiping and Xiao, Bin and Codella, Noel and Liu, Mengchen and Dai, Xiyang and Yuan, Lu and Zhang, Lei},
      title = {CvT: Introducing Convolutions to Vision Transformers},
      booktitle = {Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV)},
      year = {2021},
      pages = {22-31}
    }
  • Deep High-Resolution Representation Learning for Human Pose Estimation

    Ke Sun, Bin Xiao, Dong Liu, Jingdong Wang

    CVPR 2019

    HRNet maintains high-resolution representations throughout the network through parallel multi-resolution branches and repeated feature fusion.

    BibTeX
    @InProceedings{Sun_2019_CVPR,
      author = {Sun, Ke and Xiao, Bin and Liu, Dong and Wang, Jingdong},
      title = {Deep High-Resolution Representation Learning for Human Pose Estimation},
      booktitle = {Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)},
      year = {2019},
      pages = {5693-5703}
    }
  • Simple Baselines for Human Pose Estimation and Tracking

    Bin Xiao, Haiping Wu, Yichen Wei

    ECCV 2018

    Simple baseline methods for human pose estimation and tracking, using a deconvolutional head for keypoint prediction.

    BibTeX
    @InProceedings{Xiao_2018_ECCV,
      author = {Xiao, Bin and Wu, Haiping and Wei, Yichen},
      title = {Simple Baselines for Human Pose Estimation and Tracking},
      booktitle = {Proceedings of the European Conference on Computer Vision (ECCV)},
      year = {2018},
      pages = {466-481}
    }

Research background

  1. 2025 – present

    Post-training for reasoning, coding, and agentic models, including planning and tool use.

  2. 2024 – 2025

    Multimodal Llama post-training for visual understanding.

  3. 2024

    Led Phi-3 Vision and Phi-3.5 Vision, compact multimodal language models.

  4. 2020 – 2023

    Led the Florence-1 and Florence-2 vision foundation models.

  5. 2018 – 2021

    Developed visual representations and pose estimation methods in CvT, HRNet, and Simple Baselines.

Recognition

Selected challenge results.

  • 2019

    1st place, Look into Person Challenge, Single-Person Pose Estimation.

  • 2019

    2nd place, Object365 Challenge, Full track.

  • 2018

    1st place, PoseTrack Multi-Person Pose Tracking Challenge.

  • 2018

    2nd place, COCO Keypoint Detection Challenge.

Citation-based recognition: HRNet (CVPR 2019), CvT (ICCV 2021), and Simple Baselines (ECCV 2018) appear in Paper Digest's most-influential-paper lists (March 2026).

Contact & profiles

For professional correspondence, please connect via LinkedIn.

Scholar
Google Scholar
Location
Microsoft Superintelligence Team · Redmond, WA