Core Author
ByteDance Seed, 2026
project page /
tech
blog
Seedream 5.0 Pro is a multimodal image creation model that reasons before it draws. It
plans logical layouts for text-dense infographics, grounds point, lasso, sketch, and color signals for
pixel-level interactive editing, decomposes an image into 10+ editable transparent layers, and natively
renders text in over ten languages.
High-Resolution Visual Reasoning via Multi-Turn Grounding-Based Reinforcement
Learning
Xinyu Huang, Yuhao Dong, Wei Li, Jinming Wu, Zihao Deng, Weiwei Tian, Bo Li, Rui Feng,
Zejun Ma, Ziwei Liu
ACL 2026
arXiv /
code
MGPO lets LMMs iteratively focus on key image regions through automatic grounding, achieving superior
performance on high-resolution visual tasks without any grounding annotations.
Open-Set Image Tagging with Multi-Grained Text Supervision
Xinyu Huang, Yi-Jie Huang, Youcai Zhang, Weiwei Tian, Rui Feng, Yuejie Zhang, Yanchun Xie,
Yaqian Li, Lei Zhang
CVPR 2024, Multimodal Foundation Models Workshop
arXiv /
code
RAM++ is the next generation of RAM, which recognizes any category with high accuracy,
covering both predefined common categories and diverse open-set categories.
Recognize Anything: A Strong Image Tagging Model
Youcai Zhang*, Xinyu Huang*, Jinyu Ma*, Zhaoyang Li*, Zhaochuan Luo, Yanchun Xie, Yuzhuo
Qin, Tong Luo, Yaqian Li, Shilong Liu, Yandong Guo, Lei Zhang
CVPR 2024 Workshop on Multimodal Foundation Models
project page /
arXiv /
demo /
code
RAM is an image tagging model that recognizes any common category with high accuracy.
Tag2Text: Guiding Vision-Language Model via Image Tagging
Xinyu Huang, Youcai Zhang, Jinyu Ma, Weiwei Tian, Rui Feng, Yuejie Zhang, Yaqian Li,
Yandong Guo, Lei Zhang
ICLR 2024
project page /
arXiv /
demo /
code
Tag2Text is a vision-language model guided by tagging, which supports tagging and comprehensive
captioning simultaneously.
Increasing Text Diversity via Online Multi-Label Recognition for Vision-Language
Pre-training
Xinyu Huang, Youcai Zhang, Ying Cheng, Weiwei Tian, Ruiwei Zhao, Rui Feng, Yuejie Zhang,
Yaqian Li, Yandong Guo, Xiaobo Zhang
ACM MM 2022
arXiv /
code
IDEA provides more explicit textual supervision for visual models, including valuable tags and texts
composed of multiple tags.
Youcai Zhang*, Yuhao Cheng*, Xinyu Huang*, Fei Wen, Rui Feng, Yaqian Li, Yandong Guo
arXiv 2021
arXiv /
code
Multi-label learning with missing labels (MLML) is a challenging problem. We propose two simple yet
effective methods via robust loss design.