1. X
  2. Mercor
Log inSign up
Mercor
375 posts
Mercor profile banner
user avatar

Mercor

@mercor
Organizing human intelligence to power the AI economy.
San Francisco
mercor.com/apex
Joined April 2021
30
Following
22.5K
Followers
AffiliatesAffiliatesRepliesRepliesArticlesArticlesMediaMedia
  • Pinned
    user avatar
    Mercor
    @mercor
    Aug 12
    Grok 4.6 is in the top four across three APEX productivity benchmarks. APEX-Agents: 57.5% mean score, #4 overall APEX-Accounting: 50.9% mean score, #4 overall APEX-SWE: 56.4% Pass@1, #3 overall Congratulations to @SpaceXAI.
    Image
  • user avatar
    Mercor
    @mercor
    Aug 12
    Grok 4.6 is one of the most cost efficient frontier models we’ve ever tested on APEX.
    user avatar
    Elon Musk
    X
    @elonmusk
    Aug 12
    Grok 4.6 is now out 🚀🚀🚀 Smart, fast & amazing bang for buck!
  • user avatar
    Mercor
    @mercor
    Aug 11
    SWE-Marathon measures long-horizon development with a focus on backend tasks, but it doesn’t test whether models can build SaaS products from end to end. The @mercor research team was curious if they could, so we decided to find out. We gave eight frontier models an empty
    Image
  • user avatar
    Mercor
    @mercor
    Aug 7
    We're hiring! We have 75+ open roles across our SF, NYC, and London offices. Join us to organize human intelligence and power the AI economy. Apply today: mercor.com/careers
    Mercor - 75+ open full-time roles
  • user avatar
    Mercor
    @mercor
    Aug 5
    DeepSeek-V4-Flash was updated last week, with additional training to improve its agentic abilities. Now it outperforms their Pro model. It takes #9 overall on the APEX-Agents leaderboard with a mean score of 51.8%, just behind GLM-5.2 at 52.2%. Domain scores (mean):
    Image

Log in or sign up for X

See what’s happening and join the conversation

Continue with phone
or
Log in with username or email
Terms·Privacy·Cookies·Accessibility·Ads Info·© 2026 X Corp.
Advertisement
Advertisement