Attending #CVPR2024. Open to chat about multimodal LLMs, embodied AI, Agent and so on, pls DM me!
We will present GROUNDHOG, a multimodal large language model with pixel-level visual grounding on Thu 20 Jun 10:30 - 12 am at Arch 4A-E Poster #442.
1/ Excited to share our latest research "Grounding Visual Illusions in Language: Do Vision-Language Models Perceive Illusions Like Humans?" at #EMNLP2023 🎉 Discover how VLMs fare against tricky visual illusions👀➡️