watch this space
Who grades the graders?
LLM-as-judge on short answers is well studied. Agent judges are not: they open the files, run the code, and investigate before ruling on another agent's work.
Introducing Arbiter-Bench: 71 real agent runs that frontier agent judges get wrong.


