Overall Capabilities score is the weighted mean of each V2 benchmark's macro-mean pass rate. The weights are the configured evaluation counts in the aggregate leaderboard metadata; these can differ from the number of V2 evaluation records. Safeguards is the BioSecBench-Refusal score.
Complete-coverage model-and-harness results. Data updated 2026-09-22; artifact generated 2026-09-28T16:51:31.877Z. V2 uses clean-v1 reruns and updated evaluation sets, so scores should not be treated as a direct continuation of V1.
Source/provenance: V2 score arrays in src/App.tsx, checked against the per-evaluation files in public/data/v1/. BioSecBench-Refusal scores come from public/data/v1/refusalbench.json. Overall weights come from public/data/aggregate-leaderboard.json.
This page reports V2 agentic biology benchmark results from benchmarks.bio. V2 is a clean-v1 rerun on updated evaluation sets and currently covers fewer frontier models than V1, so the two versions are not directly comparable: the V2 overall score requires complete benchmark coverage and weights each benchmark by its configured evaluation count, while V1 equally weights whatever scores are observed. Every benchmark is built from a snapshot of real experimental data captured immediately before a target analysis step, paired with a deterministic grader that checks whether an agent recovered the actual biological result. Results are reported per model and per harness rather than per model alone, because the harness an agent runs under changes its score substantially.
Capabilities
Capabilities leaderboard across 11 benchmarks
Rank
Model
Harness
Provider
Score
1
GPT-6 Astra
Pi
OpenAI
51.3%
2
Claude Opus 5
Pi
Anthropic
49.2%
3
GPT-6 Astra
Codex
OpenAI
48.9%
4
Claude Opus 5
Claude Code
Anthropic
48.6%
5
Grok 4.6
Pi
SpaceXAI
47.8%
6
Grok 4.7
Grok Build
SpaceXAI
46.8%
7
GPT-6 Sol
Pi
OpenAI
46.8%
8
Grok 4.6
Grok Build
SpaceXAI
45.9%
9
Claude Opus 4.8
Pi
Anthropic
45.2%
10
GPT-5.6 Sol
Pi
OpenAI
44.2%
11
GPT-5.6 Sol
Codex
OpenAI
43.4%
12
Claude Opus 4.8
Claude Code
Anthropic
43.2%
13
Claude Opus 5.5
Pi
Anthropic
42.6%
14
GPT-5.5
Pi
OpenAI
42.1%
15
Claude Opus 5.5
Claude Code
Anthropic
41.8%
16
Gemini 3.7 Flash
Pi
Google
41.2%
17
Claude Sonnet 5
Pi
Anthropic
40.9%
18
Claude Sonnet 5
Claude Code
Anthropic
40.0%
19
Gemini 3.8 Flash
Pi
Google
38.5%
20
Claude Opus 4.7
Pi
Anthropic
38.5%
21
Claude Opus 4.7
Claude Code
Anthropic
38.1%
22
GPT-6 Luna
Pi
OpenAI
37.8%
23
GPT-5.6 Luna
Codex
OpenAI
34.7%
24
GPT-5.6 Luna
Pi
OpenAI
34.2%
25
Claude Sonnet 4.6
Pi
Anthropic
32.0%
26
Claude Sonnet 4.6
Claude Code
Anthropic
31.5%
Safeguards
BioSecBench-Refusal
Rank
Model
Harness
Provider
Score
1
Grok 4.7
Grok Build
SpaceXAI
62.4%
2
Gemini 3.7 Flash
Pi
Google
54.8%
3
Gemini 3.8 Flash
Pi
Google
47.4%
4
Grok 4.6
Pi
SpaceXAI
46.9%
5
Grok 4.6
Grok Build
SpaceXAI
45.6%
6
GPT-6 Astra
Pi
OpenAI
40.8%
7
Claude Opus 5
Claude Code
Anthropic
31.8%
8
GPT-6 Astra
Codex
OpenAI
25.5%
9
Claude Opus 5
Pi
Anthropic
24.5%
Included benchmarks
V2 evaluation counts can differ from the configured counts used to weight the interactive overall score.