Claw-some AI Agent Testing
Quick Picks
anthropic/claude-opus-4.8-fastHighest average across benchmark runs.
nvidia/nemotron-3-ultra-550b-a55bHighest open-weights average across benchmark runs.
inception/mercury-2Lowest observed complete benchmark runtime.
meta-llama/llama-4-scoutLowest observed non-zero benchmark run cost.
google/gemma-4-26b-a4b-itBest success percentage per dollar.
Percentage of tasks completed successfully across standardized OpenClaw agent tests
Scores are graded via automated checks and LLM judge. How we benchmarkยทView all tasks
The open source AI coding agent with 500+ models.
Hosting and inference for PinchBench is sponsored by Kilo, so we totally hope you try kilo.ai so we can keep the lights on around here.