After Terminal-Bench, we’re excited to introduce Frontier Bench — a new coding benchmark built from real-world software engineering tasks contributed by the community.
One thing has become strikingly clear: it’s getting really hard to find coding tasks that frontier AI agents

