Skip to main content

Measure everything. Publish the method.

Spider Research is the lab layer of Spider: open benchmarks and field notes on web data quality for AI. Every number links to a method; every method links to a repo.

Stealth Bench V1 80 tasks · 8 non-local entries
Spider 100% 80/80
Browser Use Cloud
81%
Anchor
74%
Onkernel
68%
Browserless
56%
Local Headful Chrome
49%
Steel
44%
Hyperbrowser
44%
Browserbase
41%
Local Headless Chrome
3%

Spider passed 80/80 anti-bot tasks on Sep 18, 2026, over loopback through the session proxy that production requests use. Spider's figure is the CDP harness run. Other providers retain the March 22, 2026 readings described in the blog as LLM-judged; they were not re-measured. Production requests use the session proxy.

Higher is better Reproduce

The record.

Each entry is a dated snapshot of a maintained measurement. Where a repo exists, the numbers reproduce from a clone.

Same tasks, same scoring, every provider.

A leaderboard is only as honest as its instrument. Every Spider benchmark holds three invariants:

01 Same tasks

Every provider runs the same 80-task list against the same pages, measuring completeness under anti-bot pressure. Each row comes from its own run, not one shared session.

02 Same scoring

A task passes when the page comes back and fails when the response matches one of the block patterns. Every provider gets the same check, and no model sits in the loop.

03 Open repo

We publish the task list and the scoring, and the runs we cite stay on this page. Clone spider-rs/benchmark and rerun the numbers yourself.

Method · spider-rs/benchmark ↗

On the bench.

  • Quality Bench v1

    Extraction fidelity, HTML against ground-truth markdown.

  • Dataset drop

    The 1,000-URL evaluation set behind the benchmark.

  • Silk structure eval

    Structure-conformance scoring for Silk output.

Judge the output yourself.

Post a URL, read what comes back. A failed request bills $0.