OpenCompass is onboard for Twitter (@X), focusing intently on the evaluation and analysis of Large Language Models and Vision-Language Models. Welcome to star our project:
😉Introduce VideoScience-Bench, a benchmark designed to evaluate undergraduate-level scientific understanding in video models.
🥰Comprises 200 carefully curated prompts spanning 14 topics and 103 concepts in physics and chemistry.
👇
hub.opencompass.org.cn/daily-benchmar…
😀A new study finds that long prompts cause a fidelity–diversity trade-off in leading T2I models: more detail but reduced diversity.
😉To evaluate this issue, the authors introduce LPD-Bench and propose PromptMoG, a training-free approach that enhances diversity by sampling
😀SpatialSky-Bench, a comprehensive benchmark specifically designed to evaluate the spatial intelligence capabilities of VLMs in UAV navigation.
😉Two categories: Environmental Perception and Scene Understanding
😉13 subcategories, including bounding boxes, color, distance,
🥰VP-Bench, a benchmark for assessing MLLMs' capability in VP perception and utilization.
😊Two-stage evaluation framework:
1.Examines models' ability to perceive VPs in natural scenes, using 30k visualized prompts spanning eight shapes and 355 attribute combinations.