pretty clear at this point what model you should use to automate labeling vision data
“why use VLMs for open-vocabulary detection? YOLO-E also supports text prompts and runs in real time.”
true. but text prompting is not the same as language understanding
using the same class names as for VLMs, YOLOE-26x scored 20.2% mAP@50 and YOLOE-11l scored 18.3%
prompts



