1. X
  2. Roboflow
Log inSign up
Roboflow
2,126 posts
Image
user avatar
Roboflow
@roboflow
Build & use computer vision models fast ✨ Get started: roboflow.com Open source datasets & models universe.roboflow.com
roboflow.com
Joined June 2019
1
Following
13.8K
Followers
RepliesRepliesMediaMedia

Log in or sign up for X

See what’s happening and join the conversation

Continue with phone
or
Log in with username or email
Terms·Privacy·Cookies·Accessibility·Ads Info·© 2026 X Corp.
  • user avatar
    Roboflow
    @roboflow
    Jul 30
    pretty clear at this point what model you should use to automate labeling vision data
    user avatar
    SkalskiP
    @skalskip92
    Jul 30
    “why use VLMs for open-vocabulary detection? YOLO-E also supports text prompts and runs in real time.” true. but text prompting is not the same as language understanding using the same class names as for VLMs, YOLOE-26x scored 20.2% mAP@50 and YOLOE-11l scored 18.3% prompts
    Image
    776
  • user avatar
    Roboflow
    @roboflow
    Jul 29
    just expanded our image reasoning evals and gemini 3.5 flash is still at the top best frontier model for visual tasks without a doubt we've benchmarked all the latest models and google takes the top 3 spots (and they are the fastest!) followed by muse spark 1.1 then fable 5
    Image
    Image
    user avatar
    SkalskiP
    @skalskip92
    Jul 29
    gemini 3.5 flash is sooooooo good at image reasoning image reasoning is when the model has to actually think about what it sees not just label objects or read text it needs to connect the facts and figure out what’s going on reasoning trace: current order top to bottom: 654,
    1.5K
  • user avatar
    Roboflow
    @roboflow
    Jul 22
    gemini 3.6 flash breakdown with example images key finding: 3.5 Flash-Lite beats Fable 5 while being ~25x cheaper gemini models continue to be the best models for vision use cases, way cheaper and faster than any other model
    user avatar
    SkalskiP
    @skalskip92
    Jul 22
    I’m both impressed and disappointed by Gemini 3.6 Flash faster, cheaper, and uses fewer tokens then Gemini 3.5 Flash. noticeably worse at object detection. often returns one general box instead of several precise ones. feels lazy. prompt: detect banana tree ↓ deep dive
    Image
    1.5K
  • user avatar
    Roboflow
    @roboflow
    Jul 21
    we benchmarked inkling on vision tasks, it's coming in at the bottom of the leaderboard. probably expected given they are leaning into fine-tuning, nothing interesting there what is interesting: it is so clear common benchmarks are saturated just a few percentage different
    Image
    Image
    user avatar
    Thinking Machines
    @thinkymachines
    Jul 15
    Today, we are introducing Inkling. Inkling reasons efficiently across text, image, and audio modalities. We are making the full weights available. thinkingmachines.ai/news/introduci… Available today for fine-tuning on Tinker. Play with it in the Inkling Playground. 🧵
    1.9K
  • user avatar
    Roboflow
    @roboflow
    Jul 2
    very good thread with sample data showing Fable 5 performance against other models for vision tasks important to see the types of tasks frontier models can/cant do will give you ideas for how to use them successfully in production
    user avatar
    SkalskiP
    @skalskip92
    Jul 2
    need a model for OCR and VQA? Fable 5 is NOT your best option Gemini 3.5 Flash is 1.3x faster, 2.9x cheaper, and more accurate Fable 5 struggles especially on tasks that require counting objects. ↓ cost comparison and interesting failure cases
    Image
    3K
  • See @roboflow's full profile

    Sign up
    Log in
Advertisement
Advertisement