Pinned
You may have heard agentic benchmarks can be gamed.
Turns out, leading solutions already do this.
With Meerkat, our framework for large-scale trace auditing, we found thousands of clear cheating instances, likely from unsupervised vibe coding.
debugml.github.io/cheating-agent…
We found widespread cheating on popular agent benchmarks, affecting 28+ submissions across 9 benchmarks and thousands of agent runs.
Surprisingly, the top 3 submissions on Terminal-Bench 2 are all cheating!
Here's what we found 🧵




