Browser Use Bench v2 Pareto frontier got completely redrawn today
> Claude Opus 5.5: 59.4
> GPT‑6 Sol medium: 66.9 (3.5x cheaper)
> GPT‑6 Luna xhigh: 57.6 (22x cheaper than Opus)
OpenAI is in its own league 🔥
All models available to try on our cloud.
[x-axis is log scale]
Open Source MiMo v2.6 models are the new Pareto frontier 🔥
> MiMo v2.6 Flash: 37.6
> Grok 4.7’s 39.9
> 57x lower recorded agent cost then Grok 4.7
> 27% cheaper then DeepSeek v4.1 Flash (low)
$0.10 vs $5.72 per task. 60 tasks each.
Open source models ❤️
Breaking: Browser Use + Jev = Ultrafast ⚡
Findings flights took 7s and cost only $0.0039 🤯
> new action space every step
> DOM state space
> small LLM fallback to type
(this video is at 1x speed btw)
Built a tiny open source browser agent. try it below ↓