Excited to share our latest work on the modality gap in multimodal LLMs.
Imagine an agent interacting with the real world as humans do. The text it encounters on signs, documents, or screens appears as pixels, not tokens. Yet MLLMs often perform worse when reading the same text
Multimodal LLMs can read text in images, but why do they often perform worse than when the same text is given as tokens? Our work studies the modality gap of models perceiving text as pixels and shows how to close it.
π arxiv.org/abs/2603.09095
π§΅π #NLProc #LLM #ComputerVision




