1. X
  2. Datacurve
Log inSign up
Datacurve
24 posts
Image
user avatar
Datacurve
@datacurve
Research and data to advance frontier models.
San Francisco
datacurve.ai
Joined February 2024
12
Following
9,781
Followers
AffiliatesAffiliatesRepliesRepliesMediaMedia

New to X?

Sign up now to get your own personalized timeline!

Create account

By signing up, you agree to the Terms of Service and Privacy Policy, including Cookie Use.

Terms·Privacy·Cookies·Accessibility·Ads Info·© 2026 X Corp.
Don't miss what's happening
People on X are the first to know.
Log inSign up
  • Pinned
    user avatar
    Datacurve
    @datacurve
    Jul 17
    Kimi K3 debuts at #3 on DeepSWE. It's the first open-weights model that delivers frontier-level performance, achieving results similar to Claude Fable and GPT-5.6 Sol.
    Image
    00:00
    649K0649K
  • user avatar
    Datacurve
    @datacurve
    Jul 17
    Replying to @datacurve
    View the full leaderboard & benchmark. →
    Image
    DeepSWE
    From deepswe.datacurve.ai
    11K011K
  • user avatar
    Datacurve
    @datacurve
    Jul 9
    GPT-5.6 tops the DeepSWE leaderboard at 73%. Sol, Terra, and Luna results are now available.
    Image
    00:00
    339K0339K
    user avatar
    Datacurve
    @datacurve
    Jul 9
    View the full family of GPT-5.6 models on our benchmark →
    Image
    DeepSWE
    From deepswe.datacurve.ai
    26K026K
  • user avatar
    Datacurve
    @datacurve
    Jul 1
    Datacurve is at @aiDotEngineer! Come by our booth in the expo hall to learn more about DeepSWE (and snag some merch). Plus, join us on Thursday @ 10:30am for our talk.
    Image
    Image
    00:00
    5.8K05.8K
  • user avatar
    Datacurve
    @datacurve
    Jun 20
    GLM 5.2 is now on DeepSWE as the top open-source model on our leaderboard. With a pass@1 score of 44% at max effort, GLM 5.2 is indisputable #1 open-source model besting Kimi K2.7 Code by 17%.
    Image
    00:00
    583K0583K
    user avatar
    Datacurve
    @datacurve
    Jun 20
    Our updated leaderboard at
    Image
    DeepSWE
    From deepswe.datacurve.ai
    15K015K
  • user avatar
    Datacurve
    @datacurve
    Jun 19
    Claude Fable 5 debuts at #1 on DeepSWE. It outscores the previous best by 3% and sets a new state-of-the-art on our long-horizon coding benchmark.
    Image
    00:00
    471K0471K
    user avatar
    Datacurve
    @datacurve
    Jun 19
    Fable 5 scores 70% pass@1 and tracks GPT-5.5 on cost-performance at the default high effort. Kimi K2.7 also joins the leaderboard with a score of 31%.
    Image
    00:00
    26K026K
    user avatar
    Datacurve
    @datacurve
    Jun 19
    See the full updated leaderboard here:
    Image
    DeepSWE
    From deepswe.datacurve.ai
    14K014K
  • Datacurve reposted
    user avatar
    Artificial Analysis
    @ArtificialAnlys
    Jun 12
    We've updated the Artificial Analysis Coding Agent Index, replacing SWE-Bench Pro with Datacurve's DeepSWE benchmark - the swap lifts Codex with GPT-5.5 (xhigh) above Claude Code with Opus 4.8 (max), while the newly released Claude Fable 5 (max) in Claude Code debuts at the top
    Image
    584K0584K
  • user avatar
    Datacurve
    @datacurve
    May 30
    Opus 4.8 is now on DeepSWE. On the default high thinking effort, it scores 6% higher than Opus 4.7 xhigh, while also lowering average cost per task.
    Image
    00:00
    982K0982K
    user avatar
    Datacurve
    @datacurve
    May 30
    Opus 4.8 delivers efficiency gains by solving tasks in fewer steps, directly reducing the total number of input tokens required per task.
    Image
    125K0125K
    user avatar
    Datacurve
    @datacurve
    May 30
    Full deep dive coming soon. Check out the full benchmark here →
    Image
    DeepSWE
    From deepswe.datacurve.ai
    26K026K
  • Datacurve reposted
    user avatar
    Matthew Berman
    Forward Future
    @MatthewBerman
    May 27
    DeepSWE reflects what I’m hearing from engineers better than any other benchmark. They took the hard path to build a good one.
    Image
    00:00
    Image
    user avatar
    Serena Ge
    Datacurve
    @serenaa_ge
    May 26
    Today we’re releasing DeepSWE, a new standard for agentic coding benchmarks. On public leaderboards, top models often look relatively close in capability. DeepSWE shows where they actually diverge, reflecting the realistic experience of developers in their day-to-day work.
    41K041K
  • Datacurve reposted
    user avatar
    Garry Tan
    Y Combinator
    @garrytan
    May 26
    This is the new standard for engineering evals
    Image
    Image
    user avatar
    Serena Ge
    Datacurve
    @serenaa_ge
    May 26
    Today we’re releasing DeepSWE, a new standard for agentic coding benchmarks. On public leaderboards, top models often look relatively close in capability. DeepSWE shows where they actually diverge, reflecting the realistic experience of developers in their day-to-day work.
    117K0117K
  • Datacurve reposted
    user avatar
    Serena Ge
    Datacurve
    @serenaa_ge
    May 26
    Today we’re releasing DeepSWE, a new standard for agentic coding benchmarks. On public leaderboards, top models often look relatively close in capability. DeepSWE shows where they actually diverge, reflecting the realistic experience of developers in their day-to-day work.
    Image
    2M02M
  • Datacurve reposted
    user avatar
    Serena Ge
    Datacurve
    @serenaa_ge
    Apr 4, 2024
    I presented today at Demo day Day 2 and @TechCrunch featured us @datacurve! Just been reading TC and listening to TC Daily Crunch since high school mornings... a surreal feeling to see us on it. Also, post-demo sadness cuz now YC is coming to an end
    Image
    Image
    31K031K
Advertisement
Advertisement