GPTNTGPTNT is a benchmark for real-time, asymmetric collaboration between multimodal agents, built on the cooperative game Keep Talking and Nobody Explodes. Against a live, unpausing clock, one agent sees the bomb but not the instructions; the other holds the instructions but never sees the bomb; neither can defuse it alone.
▸click(x=0.70, y=0.38)
▸do_nothing()
▸do_nothing()
▸out()
▸flip()
▸do_nothing()
▸do_nothing()
▸right()
▸do_nothing()
▸do_nothing()
▸right()
▸flip()
▸do_nothing()
▸do_nothing()
▸right()
▸do_nothing()
▸right()
▸do_nothing()
▸do_nothing()
▸flip()
▸do_nothing()
▸do_nothing()
▸right()
▸do_nothing()
▸right()
▸do_nothing()
▸do_nothing()
▸down()
▸do_nothing()
▸do_nothing()
▸do_nothing()
▸up()
▸do_nothing()
▸up()
▸do_nothing()
▸do_nothing()
▸down()
▸do_nothing()
▸flip()
▸do_nothing()
▸do_nothing()
▸right()
▸do_nothing()
▸right()
▸do_nothing()
▸down()
▸do_nothing()
▸do_nothing()
▸up()
▸do_nothing()
▸do_nothing()
▸do_nothing()
▸do_nothing()
▸click(x=0.70, y=0.38)
▸do_nothing()
▸click(x=0.49, y=0.52)
▸do_nothing()
▸do_nothing()
▸hold(x=0.49, y=0.52)
▸do_nothing()
▸do_nothing()
▸do_nothing()
▸release()
▸do_nothing()
▸do_nothing()
▸do_nothing()
▸do_nothing()
▸do_nothing()
▸do_nothing()
▸do_nothing()
▸do_nothing()
▸do_nothing()
▸do_nothing()
▸do_nothing()
▸do_nothing()
▸release()
▸do_nothing()
▸do_nothing()
▸do_nothing()
▸click(x=0.49, y=0.52)
▸do_nothing()
▸do_nothing()
Amit Parekh*, Sabrina McCallum*, Kareem Al-Hasan*, Malvina Nikandrou, Alessandro Suglia, Ioannis Konstas
Heriot-Watt University · University of Edinburgh
No model we test—open or closed—successfully defuses a single bomb. Nine out of ten different human pairs could solve at least one.
Models generate tokens in real time.
No rubric or LLM judge needed because the game provides the ground truth.


Our KTANE mod uses normalised (x, y) coordinates and outputs segmentation masks, so the Defuser can use coordinates or set-or-marks.
Coordinates
Set-of-marks
All twenty-three pages of rules, wiring diagrams, and symbol tables—in context from the first move.

Diagnose real collaboration by taking the manual away to see how strong the parametric knowledge is.
GPTNT runs on Keep Talking and Nobody Explodes—the same bombs, manual, and ticking clock people play against, and nothing is simplified for the models. It inherits the game’s living modding community, so as models improve we add harder modules—and eventually make them do The Centurion
One bomb with ~100 multimodal and multilingual modules.The pinnacle for any player, human or AI..



















Randomised manual solutions, pass @1
| # | Model | Interact?How did models interact with the game | Real-time(async)?Async: expert and defuser act on independent live clocks — no shared turns. | Turn-taking(sync)?Sync: expert and defuser alternate in lockstep turns. | ||||
|---|---|---|---|---|---|---|---|---|
| Missions?Full multi-module missions defused end-to-end. | Modules?Share of individual bomb modules solved across missions. | Any module?Missions where at least one module was solved before failure. | Missions | Modules | Any module | |||
| 1 | set-of-marks | 0% | 24% | 60% | 0% | 50% | 100% | |
| 2 | set-of-marks | 0% | 18% | 40% | 0% | 24% | 70% | |
| 3 | set-of-marks | 0% | 15% | 40% | 0% | 3% | 10% | |
| 4 | set-of-marks | 0% | 9% | 30% | 0% | 9% | 30% | |
| 5 | set-of-marks | 0% | 6% | 20% | 10% | 12% | 20% | |
| 6 | set-of-marks | 0% | 6% | 20% | 0% | 18% | 50% | |
| 7 | set-of-marks | 0% | 6% | 20% | 0% | 12% | 30% | |
| 8 | set-of-marks | 0% | 6% | 20% | 0% | 6% | 20% | |
| 9 | set-of-marks | 0% | 6% | 20% | 0% | 3% | 10% | |
| 10 | set-of-marks | 0% | 6% | 20% | 0% | 3% | 10% | |
| 11 | set-of-marks | 0% | 6% | 20% | 0% | 3% | 10% | |
| 12 | set-of-marks | 0% | 3% | 10% | 0% | 3% | 10% | |
| 13 | set-of-marks | 0% | 3% | 10% | 0% | 3% | 10% | |
| 14 | set-of-marks | 0% | 0% | 0% | 0% | 6% | 20% | |
| 15 | set-of-marks | 0% | 0% | 0% | 0% | 3% | 10% | |
| 16 | set-of-marks | 0% | 0% | 0% | 0% | 3% | 10% | |
| — | Human players | — | 25% | 60% | 93% | — | — | — |
Ran GPTNT on your own model?
Submit your run and we'll add it to the board. New models and protocols welcome.
Results · n=10
Mean tokens per mission
Run 16 Sep 2026 · GPTNT v2.1.1
Defuser + Expert
GPT-6 AstraOpenAI
Results · n=10
Mean tokens per mission
Run 16 Sep 2026 · GPTNT v2.1.1
Defuser + Expert
GPT-6 AstraOpenAI
Results · n=10
Mean tokens per mission
Run 08 Sep 2026 · GPTNT v2.1.1
Defuser + Expert
GPT-5.6 TerraOpenAI
Results · n=10
Mean tokens per mission
Run 08 Sep 2026 · GPTNT v2.1.1
Defuser + Expert
GPT-5.6 TerraOpenAI
Results · n=10
Mean tokens per mission
Run 08 Sep 2026 · GPTNT v2.1.1
Defuser + Expert
Inkling-SmallThinking Machines Lab
Results · n=10
Mean tokens per mission
Run 08 Sep 2026 · GPTNT v2.1.1
Defuser + Expert
Inkling-SmallThinking Machines Lab
Results · n=10
Mean tokens per mission
Run 21 Sep 2026 · GPTNT v2.1.1
Defuser + Expert
Qwen3.8 Flash Next (180B)Qwen
Results · n=10
Mean tokens per mission
Run 21 Sep 2026 · GPTNT v2.1.1
Defuser + Expert
Qwen3.8 Flash Next (180B)Qwen
Results · n=10
Mean tokens per mission
Run 08 Sep 2026 · GPTNT v2.1.1
Defuser + Expert
Claude Sonnet 5Anthropic
Results · n=10
Mean tokens per mission
Run 08 Sep 2026 · GPTNT v2.1.1
Defuser + Expert
Claude Sonnet 5Anthropic
Results · n=10
Mean tokens per mission
Run 11 Sep 2026 · GPTNT v2.1.1
Defuser + Expert
Qwen3.8 (27B)Qwen
Results · n=10
Mean tokens per mission
Run 11 Sep 2026 · GPTNT v2.1.1
Defuser + Expert
Qwen3.8 (27B)Qwen
Results · n=10
Mean tokens per mission
Run 08 Sep 2026 · GPTNT v2.1.1
Defuser + Expert
Claude Sonnet 4.6Anthropic
Results · n=10
Mean tokens per mission
Run 08 Sep 2026 · GPTNT v2.1.1
Defuser + Expert
Claude Sonnet 4.6Anthropic
Results · n=10
Mean tokens per mission
Run 10 Sep 2026 · GPTNT v2.1.1
Defuser + Expert
Qwen3.6 (35B MoE)Qwen
Results · n=10
Mean tokens per mission
Run 10 Sep 2026 · GPTNT v2.1.1
Defuser + Expert
Qwen3.6 (35B MoE)Qwen
Results · n=10
Mean tokens per mission
Run 08 Sep 2026 · GPTNT v2.1.1
Defuser + Expert
GPT-5.2OpenAI
Results · n=10
Mean tokens per mission
Run 08 Sep 2026 · GPTNT v2.1.1
Defuser + Expert
GPT-5.2OpenAI
Results · n=10
Mean tokens per mission
Run 10 Sep 2026 · GPTNT v2.1.1
Defuser + Expert
Gemma 4 (26B MoE)Google DeepMind
Results · n=10
Mean tokens per mission
Run 10 Sep 2026 · GPTNT v2.1.1
Defuser + Expert
Gemma 4 (26B MoE)Google DeepMind
Results · n=10
Mean tokens per mission
Run 17 Sep 2026 · GPTNT v2.1.1
Defuser + Expert
Mistral Small 4 (119B-A6.5B)Mistral AI
Results · n=10
Mean tokens per mission
Run 17 Sep 2026 · GPTNT v2.1.1
Defuser + Expert
Mistral Small 4 (119B-A6.5B)Mistral AI
Results · n=10
Mean tokens per mission
Run 16 Sep 2026 · GPTNT v2.1.1
Defuser + Expert
Nemotron 3 Nano Omni (30B-A3B)NVIDIA
Results · n=10
Mean tokens per mission
Run 16 Sep 2026 · GPTNT v2.1.1
Defuser + Expert
Nemotron 3 Nano Omni (30B-A3B)NVIDIA
Results · n=10
Mean tokens per mission
Run 08 Sep 2026 · GPTNT v2.1.1
Defuser + Expert
Qwen3.5 (27B)Qwen
Results · n=10
Mean tokens per mission
Run 08 Sep 2026 · GPTNT v2.1.1
Defuser + Expert
Qwen3.5 (27B)Qwen
Results · n=10
Mean tokens per mission
Run 17 Sep 2026 · GPTNT v2.1.1
Defuser + Expert
Step 3.7 Flash (198B-A11B)StepFun
Results · n=10
Mean tokens per mission
Run 17 Sep 2026 · GPTNT v2.1.1
Defuser + Expert
Step 3.7 Flash (198B-A11B)StepFun
Results · n=10
Mean tokens per mission
Run 10 Sep 2026 · GPTNT v2.1.1
Defuser + Expert
Gemma 4 (31B)Google DeepMind
Results · n=10
Mean tokens per mission
Run 10 Sep 2026 · GPTNT v2.1.1
Defuser + Expert
Gemma 4 (31B)Google DeepMind
Results · n=10
Mean tokens per mission
Run 10 Sep 2026 · GPTNT v2.1.1
Defuser + Expert
Qwen3.6 (27B)Qwen
Results · n=10
Mean tokens per mission
Run 10 Sep 2026 · GPTNT v2.1.1
Defuser + Expert
Qwen3.6 (27B)Qwen
Latest update · 29 Sep 2026
The leaderboard now reports results from 16 models on GPTNT’s randomised manual solutions, including first results for Claude Sonnet 5, GPT-5.6 Terra, GPT-6 Astra, Gemma 4, Inkling-Small, Mistral Small 4, Nemotron 3 Nano Omni, Qwen3.6, Qwen3.8, and Step 3.7 Flash. The earlier original-solution leaderboard remains available as a historical comparison.
We’re delighted to share that GPTNT has been accepted by Transactions on Machine Learning Research (TMLR). 🎊
GPTNT now uses randomised manual solutions. Each benchmark suite selects a rule seed that changes the solution logic for supported modules, and GPTNT compiles a matching manual for the expert agent. In other words, the bomb’s rules and the expert’s instructions still line up—but they are no longer simply the familiar, original KTANE solutions that a model might have encountered in training. This gives us a more direct protection against memorisation, so we no longer need separate contamination checks. Those checks have now been retired.
Of actual games played by models we tested
Defuser viewAsync replays coming soon
Parekh, McCallum, AlHasan, Nikandrou, Suglia, Konstas
Transactions on Machine Learning Research · 2026