Outerport@outerportJun 30, 2025Structured Attention Matters to Multimodal LLMs in Document Understanding arxiv.org/abs/2506.216004936
Outerport@outerportJun 30, 2025doctors hate himTowaki Takikawa / 瀧川永遠希@yongyuanxiJun 30, 2025Adding horizontal lines to images improves VLM (vision language model) performance of tasks like counting, visual search, spatial understating, scene understanding, and more41.2K
Towaki Takikawa / 瀧川永遠希@yongyuanxiJun 30, 2025Adding horizontal lines to images improves VLM (vision language model) performance of tasks like counting, visual search, spatial understating, scene understanding, and more