Pinned
New evals from Turing on Terminal-Bench 3.0
Capable agents can pass their own checks and still be confidently wrong. We analyzed 3,600 Terminal-Bench 3.0 trials to understand why.
The pattern is surprisingly consistent: agents often solve the wrong problem, verify against the

