Last night we hosted our 3rd @scale_AI Research Meetup: Benchmarking Coding Agents at our NYC office!
Coding agents are improving faster than our ability to measure them. That gap was the thread running through the whole evening: how do we build evals that actually tell us what
welcome to the lab.
from the researchers at @scale_AI
Joined October 2025
- AI safety works best when it's collaborative. Proud that our red team had early access to @thinkymachines Inkling to evaluate misuse risks and policy compliance before launch. Independent evaluation before release helps developers make evidence-based decisions and strengthensReleasing weights indiscriminately isn't safe. Neither is keeping capable models inside a few labs. We think there's a path between them. We haven't mapped all of it. Our new post covers the part we can see: how we assessed Inkling, and why access should widen in stages.
- Congrats to @thinkymachines on the release of Inkling-small! A smaller variant of Inkling, now live on our AudioMultiChallenge and MCP Atlas leaderboards. Inkling-small is tied for🥇on AudioMultiChallenge, scoring about the same as the larger Inkling despite the size difference.
- Calling the NYC AI research community 📣 Our third Research Meetup is almost here, and this time we're diving into benchmarking for coding agents. Join us at the @scale_AI NYC office to hear from researchers building at the frontier — plus food, drinks, and swag. Speakers:
- Today we're releasing the complete EnigmaEval dataset, our benchmark built from sophisticated puzzle-hunt problems in partnership with @CAIS. We believe open access to EnigmaEval will help push reasoning and evaluation research forward. So far, Claude Fable 5 and GPT-5.6 Sol


