I have been using LLM as verifier from early 2026. I have been using and integrating this github in the last few months ad well, also making different variants of it with specific verification models.
All my agents / workflows are using LLM as verifier. Any single task uses
Scaling self-verification with DeepSeek V4 Flash beats Claude Fable 5 on Terminal-Bench 2.1, while being 11x cheaper 💰
As open-source models become more capable, they can now generate large numbers of high-quality candidate solutions and verify their own outputs at very low







