Log inSign up
Mercor
397 posts
Mercor profile banner
@mercor

Mercor

@mercor
Organizing human intelligence to power the AI economy.
San Francisco
mercor.com/apex
Joined April 2021
30
Following
23.4K
Followers
RepliesRepliesRepostsRepostsMediaMediaArticlesArticles

Log in or sign up for X

See what’s happening and join the conversation

Continue with phone
or
Log in with username or email
Terms·Privacy·Cookies·Accessibility·Ads Info·© 2026 X Corp.
  • Pinned
    @mercor
    Mercor
    @mercor
    Aug 26
    Today we're introducing the Mercor Research Fellowship. We're funding a small group of people to build a new APEX benchmark, the definitive measure of frontier AI on professional-quality work. Applications are now open: mercor.com/careers/?ashby…
    Image
    8
  • @mercor
    Mercor
    @mercor
    Sep 5
    GPT-6 Astra passes more tasks on APEX-Accounting than any other model. 13.1% Pass@1 (#1) 60.0% mean score (#2) Pass@1 is the proportion of tasks that a model scores 100% at least once across four attempts. Astra passes 56% more tasks than GPT-5.6 Sol and 12% more tasks than
    Image
    1
  • @mercor
    Mercor
    @mercor
    Sep 3
    GPT-6 Astra debuts on APEX-Agents at #1 on the leaderboard. 🥇 62.4% Mean score (#1) 🥈 46.7% Pass@1 (#2) It leads Fable 5.1 by 0.4 points on mean score and trails it by 0.8 on Pass@1. APEX-Agents runs long-horizon tasks written by bankers, consultants, and corporate lawyers
    Image
    7
  • @mercor
    Mercor
    @mercor
    Sep 1
    Data is the most important ingredient in post-training. The Mercor Research team focuses on making every hour of expert work yield the most model improvement, through better learning algorithms for knowledge work and automated, domain-specific post-training. We're also
    Training frontier knowledge work agents: A 397B RL training guide with SkyRL
    Training frontier knowledge work agents: A 397B RL training guide with SkyRL | Mercor Blog
    From mercor.com
    3
  • @mercor
    Mercor
    @mercor
    Sep 1
    Fable 5.1 debuts at #2 on the APEX-SWE leaderboard, within the confidence band of first place. It’s also the new leader for Integration tasks. Overall: 63.6% (#2) Integration: 68.1% (#1) Observability: 59.0% (#2) Congratulations to @AnthropicAI @claudeai.
    Image
    3
Advertisement
Advertisement