1. X
  2. Surge AI
Log inSign up
Surge AI
710 posts
Image
user avatar
Surge AI
@HelloSurgeAI
Our mission is to raise AGI with the richness of humanity — curious, witty, imaginative, and full of breathtaking brilliance.
surgehq.ai
Joined June 2020
141
Following
8,701
Followers
RepliesRepliesMediaMedia
  • user avatar
    Surge AI
    @HelloSurgeAI
    Aug 6
    The easy coding problems are solved. The hard ones are the whole job now. We're hiring senior engineers to find where frontier models break and set the bar they're trained against. Remote, $100-150+/hr, no ML background needed.
    surgehq.ai
    Senior Software Engineer | Surge AI
    Make $200-300k+/year building agentic coding benchmarks and RL environments. Fully remote, $100-$150+/hour, set your own hours.
  • user avatar
    Surge AI
    @HelloSurgeAI
    Aug 5
    We post-trained a model on office work. It also improved at coding. New paper from our research team! The training run: Office work RL environments (spreadsheets, documents, web research, planning). Zero coding tasks. The result: +5.8pp on SWE-Bench Pro. What transferred: A
  • user avatar
    Surge AI
    @HelloSurgeAI
    Aug 3
    Most engineers will use AI. A few will get to shape it. We're hiring senior engineers to build the coding benchmarks frontier labs train on. Find where the best models break, set the bar they're measured against. Remote, contract, $100-150+/hr, no ML background needed.
    surgehq.ai
    Senior Software Engineer | Surge AI
    Make $200-300k+/year building agentic coding benchmarks and RL environments. Fully remote, $100-$150+/hour, set your own hours.
  • user avatar
    Surge AI
    @HelloSurgeAI
    Jul 31
    Frontier Data, Off the Shelf. The training data, evals, and RL environments we build for frontier labs, now available directly. ⌨️ Agentic coding: 35.2 → 47.2 on Terminal-Bench 2.0 🔧 Enterprise agents: 24.2 → 33.8 on Toolathlon 📐 STEM reasoning: 29.1 → 38.4 on
    Image
  • user avatar
    Surge AI
    @HelloSurgeAI
    Jul 27
    Our Chartography benchmark measures whether models can read professional charts. @AnthropicAI’s contribution to science: what if it squinted? Nice job Fable.
    user avatar
    ClaudeDevs
    Anthropic
    @ClaudeDevs
    Jul 23
    Replying to @ClaudeDevs
    On the Chartography benchmark (100 questions over dense real-world charts), Fable 5's accuracy goes from 29% to 73% with a zoom tool, and Sonnet 5’s from 13% to 44%. See the cookbook to learn more: github.com/anthropics/cla…

Log in or sign up for X

See what’s happening and join the conversation

Continue with phone
or
Log in with username or email
Terms·Privacy·Cookies·Accessibility·Ads Info·© 2026 X Corp.
Advertisement
Advertisement