At depthfirst, we evaluated Opus 5.5 on our vulnerability detection eval.
From my investigation, it's more selective and doesn't search as hard as OpenAI models.
It is definitely an improvement on Opus-5, being more token efficient while having overall better performance!
Research Engineer @ DepthFirst reducing cyber risk. Prev AI Safety Fellow @ Redwood catching misaligned models. Previously exited an AI startup.





