benchmarks.bio

V1 results

V1 cross-benchmark results

Overall score is the equally weighted mean of observed benchmark scores.

Data updated 2026-09-20; artifact generated 2026-09-21T21:43:49.599Z. Missing scores are excluded from a model's mean and are never treated as zero.

benchmarks.bio publishes agentic AI benchmark results across three domains: omics, therapeutics, and biosecurity. Every benchmark is built from a snapshot of real experimental data captured immediately before a target analysis step, paired with a deterministic grader that checks whether an agent recovered the actual biological result rather than whether it produced plausible-looking output. Results are reported per model and per harness rather than per model alone, because the harness an agent runs under changes its score substantially. The overall score in the table below is the equally weighted mean of a configuration's observed scores across the 11 capability benchmarks; a missing benchmark score is excluded from that mean and is never counted as zero. V1 and V2 cover different evaluation sets and should not be read as a single time series.

Capabilities leaderboard across 11 benchmarks
RankModelHarnessProviderOverall score
1GPT-6 AstraPiOpenAI51.3%
2GPT-6 AstraCodexOpenAI51.1%
3Claude Opus 5Claude CodeAnthropic48.6%
4Claude Opus 5PiAnthropic48.2%
5Grok 4.6Grok BuildSpaceXAI48.2%
6GPT-5.6 SolPiOpenAI47.5%
7GPT-5.6 SolCodexOpenAI45.9%
8DeepSeek V4.1 FlashPiDeepSeek45.0%
9Grok 4.6PiSpaceXAI44.8%
10Claude Opus 4.8PiAnthropic44.8%
11Claude Opus 4.8Claude CodeAnthropic42.8%
12GPT-5.6 TerraPiOpenAI42.4%
13Gemini 3.5 FlashPiGoogle41.7%
14GPT-5.6 TerraCodexOpenAI41.5%
15Grok 4.5PiSpaceXAI41.3%
16GPT-5.5PiOpenAI40.1%
17Claude Opus 4.7Claude CodeAnthropic39.8%
18Claude Sonnet 5PiAnthropic39.5%
19Claude Sonnet 5Claude CodeAnthropic39.1%
20GPT-5.5CodexOpenAI38.7%
21Kimi K3PiMoonshot38.1%
22GPT-5.6 LunaPiOpenAI33.8%
23GPT-5.6 LunaCodexOpenAI32.2%

Included benchmarks

  • SpatialBench — Omics; published score: Verified subset (115 evaluations); Full benchmark: 159 evaluations
  • SpatialBench-Long — Omics; published score: Verified subset (22 evaluations); Full benchmark: 24 evaluations
  • scBench — Omics; published score: Full benchmark (195 evaluations)
  • scBench-Long — Omics; published score: Verified merged subset (22 evaluations); Original full release: 21 evaluations
  • EpiBench — Omics; published score: Full benchmark (106 evaluations)
  • VariantBench — Omics; published score: Full benchmark (118 evaluations)
  • TxBench-Antibody-Discovery — Therapeutics; published score: Full benchmark (100 evaluations)
  • TxBench-Preclinical-Pharmacology — Therapeutics; published score: Full benchmark (100 evaluations)
  • TxBench-Oligo-Discovery — Therapeutics; published score: Full benchmark (113 evaluations)
  • BioSecBench-Surveillance — Biosecurity; published score: Full benchmark (102 evaluations)
  • BioSecBench-Refusal — Biosecurity; published score: Full benchmark (107 evaluations)
  • BioSecBench-Function — Biosecurity; published score: Full benchmark (111 evaluations)