Agents can open PRs faster than any team can review them.
There’s a tool to protect your judgement from getting spent on the wrong work.
It’s called CodeRabbit Triage.
Beyond the benchmarks, what was Opus 5.5 like to work with?
Hendrik and Gowtham discuss an overnight coding run, the token tradeoffs, and why higher reasoning effort didn’t consistently catch more bugs.
We ran Opus 5.5 through CodeRabbit's review pipeline.
> On 80 known bug patterns it caught 51 vs 49 for our production mix.
> On 13 harder cases, 10 vs 5.
> It found a retry-count race in Cal.com that production missed.
The catch is ~50% more tokens, and 9 bugs our baseline
We ran Opus 5.5 through CodeRabbit's review pipeline.
> On 80 known bug patterns it caught 51 vs 49 for our production mix.
> On 13 harder cases, 10 vs 5.
> It found a retry-count race in Cal.com that production missed.
The catch is ~50% more tokens, and 9 bugs our baseline
Introducing Claude Opus 5.5, the first model in our new Claude 5.5 family.
It performs at the level of Claude Fable 5.1 for most tasks, and costs 40% less to run than Opus 5.