Pinned
On ITBench, Qwen3.5-27B with an evolved harness beats GPT-5.5 on the baseline harness by +19.2 points.
More capable models do not always make better enterprise agents.
With the right harness, a smaller open-weight model can outperform a larger frontier model running the default



