Log inSign up
Handshake
3,121 posts
Handshake profile banner
@joinHandshake

Handshake

@joinHandshake
Building the future workforce of the AI economy 🤝
United States
joinhandshake.com
Joined March 2014
352
Following
11.4K
Followers
RepliesRepliesRepostsRepostsMediaMedia

Log in or sign up for X

See what’s happening and join the conversation

Continue with phone
or
Log in with username or email
Terms·Privacy·Cookies·Accessibility·Ads Info·© 2026 X Corp.
  • Pinned
    @joinHandshake
    Handshake
    @joinHandshake
    Jun 10, 2025
    Introducing Handshake AI—the most ambitious chapter in our story. We leverage the scale of the largest early career network to source, train, and manage domain experts who test and challenge frontier models to failure for the top AI labs.
    Image
    00:00
    38
  • @joinHandshake
    Handshake
    @joinHandshake
    Aug 25
    Today we're launching ATLAS Visual Life Sciences (VIALS), a benchmark testing whether AI can interpret the images life scientists make decisions from. 10 frontier models. 161 tasks from real biotech and pharma workflows. Best score so far: 26.5% — a clear, measurable gap to
    Image
    15
  • @joinHandshake
    Handshake
    @joinHandshake
    Aug 25
    In-depth conversation on evals, agents and what enterprise AI adoption takes with @GarrettLord and @Sbhaiwala03
    @MTSlive
    MTS
    @MTSlive
    Aug 24
    Handshake AI CSO @Sbhaiwala03 reveals how they build simulated white-collar work environments to train frontier AI agents: "An environment consists of a few elements. The first is the software tools the agent needs access to. If I'm an investment banker, I have access to Excel,
    Image
    00:00
    2
  • @joinHandshake
    Handshake
    @joinHandshake
    Aug 11
    Landed on Pavlov’s List! Exciting work from our research team @guzmanhe, @awws0me, @vaibhav4595, @andreas_plesner, @anishathalye, and Yi Liu.
    @chrisbarber
    Chris Barber
    @chrisbarber
    Aug 3
    i made some updates to pavlov’s list and am pleased with them! thanks to @xeophon @ybenpan @phoebeyao @amit05prakash @kate_shapova @timshi_ai @kevinhou22 @madiator @dlbydq @daljeet_v @ingmariusX @davidstutz92 & Dylan Rogers for feedback lately (pavlov's = list of rl env cos)
    Image
    5
  • @joinHandshake
    Handshake
    @joinHandshake
    Jul 25
    Our team is proud to have contributed to Frontier-Bench! As agents take on more ambitious work, our benchmarks must become more ambitious too. Exciting to see Anthropic's Opus 5 release today already substantially improve upon Fable 5 from 33% -> 43% on this benchmark.
    @ryan_marten
    Ryan Marten
    @ryan_marten
    Jul 23
    We’re releasing Frontier-Bench: a benchmark that measures and evolves with the frontier of agent work. Built by the team behind Terminal-Bench and Harbor, Frontier-Bench is an on-going community effort. Frontier-Bench v0.1 contains 74 tasks on which the best agents score ~34%
    Image
    5
Advertisement
Advertisement