Pinned
We got 46% fewer errors than the single best LLM across the 16 most used benchmarks (TerminalBench, LiveCodeBench, etc).
Here's how that's possible and what each model can achieve when used optimally (every benchmarks misses the majority of model capabilities) 👇
Interactive


