In this work, we revisit how automatic harness evolution should be evaluated.
Existing automatic harness evolution methods often search over harnesses using feedback from benchmark tasks and then report final performance on the same benchmark. This makes it difficult to tell
Automatic harness evolution appears to be a promising path toward AI self-improvement, but we find that its gains still largely come from repeated sampling and show limited generalization.
Blog post: yikee.github.io/harnessevoluti…
Code: github.com/rethinking-har…





