Introducing Handshake AI—the most ambitious chapter in our story. We leverage the scale of the largest early career network to source, train, and manage domain experts who test and challenge frontier models to failure for the top AI labs.
Our team is proud to have contributed to Frontier-Bench!
As agents take on more ambitious work, our benchmarks must become more ambitious too.
Exciting to see Anthropic's Opus 5 release today already substantially improve upon Fable 5 from 33% -> 43% on this benchmark.
We’re releasing Frontier-Bench: a benchmark that measures and evolves with the frontier of agent work.
Built by the team behind Terminal-Bench and Harbor, Frontier-Bench is an on-going community effort.
Frontier-Bench v0.1 contains 74 tasks on which the best agents score ~34%
We came to Seoul for #ICML. We stayed for the tteokbokki!
Last night we hosted a group of researchers at Gwangjang Market, one of Seoul's oldest night markets, for a private food tour. We ate our way through the stalls, and yes, it lived up to every bit of the hype.
But the
Day 1 @icmlconf is officially in the books! 🇰🇷
The Handshake AI booth was buzzing all day long with great conversations, sharp questions, and even sharper minds. Thank you to everyone who stopped by!
Haven't made it by yet? We're just getting started. Swing by booth B600 on
We worked with parents and professionals in child-protection and clinical psychology to test 7 frontier AI models on child safety scenarios that go beyond the frequent focus on explicit content. The parents catch what standards evaluations don't, and the professionals bring real,
AI models pose serious child-safety risks. While many model developers evaluate for explicit abuse material, other child-safety failures begin upstream: when a model helps an adult manipulate, impersonate, profile, or isolate a minor; or when it deepens a child’s emotional