Grok 4.6 is in the top four across three APEX productivity benchmarks.
APEX-Agents: 57.5% mean score, #4 overall
APEX-Accounting: 50.9% mean score, #4 overall
APEX-SWE: 56.4% Pass@1, #3 overall
Congratulations to @SpaceXAI.
SWE-Marathon measures long-horizon development with a focus on backend tasks, but it doesn’t test whether models can build SaaS products from end to end.
The @mercor research team was curious if they could, so we decided to find out.
We gave eight frontier models an empty
We're hiring! We have 75+ open roles across our SF, NYC, and London offices.
Join us to organize human intelligence and power the AI economy.
Apply today: mercor.com/careers
DeepSeek-V4-Flash was updated last week, with additional training to improve its agentic abilities. Now it outperforms their Pro model.
It takes #9 overall on the APEX-Agents leaderboard with a mean score of 51.8%, just behind GLM-5.2 at 52.2%.
Domain scores (mean):