Introducing Handshake AI—the most ambitious chapter in our story. We leverage the scale of the largest early career network to source, train, and manage domain experts who test and challenge frontier models to failure for the top AI labs.
Today we're launching ATLAS Visual Life Sciences (VIALS), a benchmark testing whether AI can interpret the images life scientists make decisions from.
10 frontier models. 161 tasks from real biotech and pharma workflows. Best score so far: 26.5% — a clear, measurable gap to
Handshake AI CSO @Sbhaiwala03 reveals how they build simulated white-collar work environments to train frontier AI agents:
"An environment consists of a few elements. The first is the software tools the agent needs access to. If I'm an investment banker, I have access to Excel,
Our team is proud to have contributed to Frontier-Bench!
As agents take on more ambitious work, our benchmarks must become more ambitious too.
Exciting to see Anthropic's Opus 5 release today already substantially improve upon Fable 5 from 33% -> 43% on this benchmark.
We’re releasing Frontier-Bench: a benchmark that measures and evolves with the frontier of agent work.
Built by the team behind Terminal-Bench and Harbor, Frontier-Bench is an on-going community effort.
Frontier-Bench v0.1 contains 74 tasks on which the best agents score ~34%