Introducing our Browser Agent Evals.
We benchmark frontier and open models using various harnesses on computer-use tasks.
Our evals are fully open source and reproducible via our Evals CLI.
GPT-6 Luna is a step function improvement from it's predecessor scoring 11% higher in accuracy on our browser agent tasks.
GPT-6 Sol scores 2% better in accuracy, but costs 20x more per task than Luna.
GPT-6 Sol and Luna just landed in Astra’s orbit.
Both launch today with API prices 50% lower than GPT-5.6.
Build with Sol. Scale with Luna. To production and beyond.
Stagehand can use Stripe's new WebMCP tools in checkout pages.
Run page.tools() to see what WebMCP connectors a page exposes. With Stripe, you can use Stagehand to add items to cart and complete the checkout with native tools.
Hi internet! We went ahead and upgraded all @stripe hosted checkout pages for 7.8M businesses—0.45% of the world’s GDP—with WebMCP so agents can more efficiently buy. Our evals show this reduces checkout latency by 39% and token usage by 42%.
Learn more: stripe.dev/blog/how-strip…
During early testing on our Browser Agent Evals, Opus 5.5 completely crushed it's predecessor Opus 5, but also Fable 5.1 & 5 in accuracy, speed, and, cost.
Compared to Opus 5 in the Claude Code harness it's roughly 5x cheaper per task, and 2x faster.
Introducing Claude Opus 5.5, the first model in our new Claude 5.5 family.
It performs at the level of Claude Fable 5.1 for most tasks, and costs 40% less to run than Opus 5.
it seems like agentic commerce is finally solved with stripe link. and we're excited to partner with stripe to enable all agents to transact online.
your agent can now natively make payments on the web using link cli + browserbase.
one shot prompt: docs.browserbase.com/integrations/s…