dig.bench
Discovering unknown rules in text-based games
Leaderboard
All games are human beatable, on first attempt, as validated on external human testers. Win rate: each game's wins are averaged over its runs, then those are averaged across the tier's ten games.
About
dig.bench is a benchmark of scientific discovery.
Each of its 70 games measures whether an agent can experiment to discover that game's own unknown rules. Every game is text-based, which puts it in the natural domain of language models: no visual confounds stand between a model and the discovery, so what dig.bench tests is discovery alone. Humans and frontier models play the same games with access to the same information, and progress is scored by whether the game can be beaten within a limited number of steps.
The games come in 7 tiers, depending on their difficulty. No game is easy and they all require effortful play, but humans can make the discoveries necessary to solve even our hardest games, while the best models struggle to beat games in the top tier.
- Scale
- 70 new interactive games (21 publicly released).
- What qualities of models do we test
- To beat each game an agent must discover the unknown rules and apply them to solve challenges.
- Evaluation
- Humans and frontier models play through the same interface: identical game states, identical action sets, identical step budgets.

Play
21 of the 70 games are public. Tiers get increasingly harder for models (1 = easiest, 7 = hardest).
Reproduce it
Run any model against the games through the SDK or API.
Join us
Join our community DisCo, where you can track your progress on these puzzles and hang out with like-minded folk.
Citation
DiG-bench: Discovery in Games
@misc{battleday2026dig,
title={DiG-bench: Discovery in Games},
author={Ruairidh M. Battleday and Kai Sandbrink and Jimi Cullen-Drohan and Zihan Yan and Timothy Muller and Clare Maguire and Ales Kubicek and Fraser Greenlee-Scott and Sukrit Sumant and Tri Dao and Jürgen Schmidhuber and Michal Valko and Joshua Tenenbaum and Thomas L. Griffiths and Zeb Kurth-Nelson and James C.R. Whittington},
year={2026},
eprint={2608.12593},
archivePrefix={arXiv},
primaryClass={cs.AI},
url={https://arxiv.org/abs/2608.12593},
}