Log inSign up
Log inSign up
Handshake
3,144 posts
Handshake profile banner
@joinHandshake

Handshake

@joinHandshake
Building the future workforce of the AI economy 🤝
United States
joinhandshake.com
Joined March 2014
351 Following
11.9K Followers
RepliesRepliesRepostsRepostsMediaMedia

Log in or sign up for X

See what’s happening and join the conversation

Continue with phone
or
Log in with username or email
Terms·Privacy·Cookies·Accessibility·Ads Info·© 2026 X Corp.
  • Pinned
    @joinHandshake
    Handshake
    @joinHandshake
    Jun 10, 2025
    Introducing Handshake AI—the most ambitious chapter in our story. We leverage the scale of the largest early career network to source, train, and manage domain experts who test and challenge frontier models to failure for the top AI labs.
    Image
    00:00
    42
  • @joinHandshake
    Handshake
    @joinHandshake
    Oct 10
    What a week at @COLM_conf in SF! We loved meeting the research community, hearing about the work people are doing, and sharing more about our ATLAS benchmark collection at the Handshake Al booth. Also, a huge thank you to our friends at @cerebras for cohosting Cafe Compute with
    Image
    Image
    Image
    Image
  • @joinHandshake
    Handshake
    @joinHandshake
    Oct 8
    “The model’s capabilities are pretty spiky.” — Jonas Mueller
    @jomulr
    Jonas Mueller @ COLM 2026
    @jomulr
    Oct 8
    The ultimate benchmark for new models is how useful professionals find them to get stuff done. Explaining benchmaxxing on The Information TV this morning, and why there’s no substitute for actual feedback from scientists/doctors using the model to solve their problems.
    Image
    00:00
    2
  • @joinHandshake
    Handshake
    @joinHandshake
    Oct 8
    The best GRE prep used to be a budget question. This benchmark says it doesn't have to be.
    @nishantXranka
    Nishant
    @nishantXranka
    Oct 8
    @joinHandshake ran a controlled study with 2800+ students comparing the impact of AI with human tutor. Two interesting findings: (1) AI and human produced similar learning gains, and (2) AI is 70x (Opus) to ~900x (Gemma) cheaper. See: ATLAS-Education joinhandshake.com/research/bench…
    1
  • @joinHandshake
    Handshake
    @joinHandshake
    Oct 6
    Agent benchmarks should be grounded in how real people actually interact. This often requires user simulation, but current models often lack fidelity. After months of development, our research team has published CUE, a simulator that outperforms state-of-the-art approaches on
    @anjali_ruban
    Anjali Kantharuban
    @anjali_ruban
    Oct 6
    ⚠️ User simulators are increasingly used to evaluate AI agents in multi-turn interaction. But we find that looking like real users ≠ reproducing real evaluation outcomes. 📢 We introduce Calibrated User Embeddings (CUE) to bridge this gap across LLM simulators. 🧵 1/8
    Image
Advertisement
Advertisement