1. X
  2. Scale Labs
Log inSign up
Scale Labs
160 posts
Image
user avatar
Scale Labs
@ScaleAILabs
welcome to the lab. from the researchers at @scale_AI
labs.scale.com
Joined October 2025
112
Following
2,363
Followers
RepliesRepliesMediaMedia
  • user avatar
    Scale Labs
    @ScaleAILabs
    Aug 4
    AI safety works best when it's collaborative. Proud that our red team had early access to @thinkymachines Inkling to evaluate misuse risks and policy compliance before launch. Independent evaluation before release helps developers make evidence-based decisions and strengthens
    user avatar
    Thinking Machines
    @thinkymachines
    Jul 31
    Releasing weights indiscriminately isn't safe. Neither is keeping capable models inside a few labs. We think there's a path between them. We haven't mapped all of it. Our new post covers the part we can see: how we assessed Inkling, and why access should widen in stages.
  • user avatar
    Scale Labs
    @ScaleAILabs
    Jul 30
    Congrats to @thinkymachines on the release of Inkling-small! A smaller variant of Inkling, now live on our AudioMultiChallenge and MCP Atlas leaderboards. Inkling-small is tied for🥇on AudioMultiChallenge, scoring about the same as the larger Inkling despite the size difference.
    Image
  • user avatar
    Scale Labs
    @ScaleAILabs
    Jul 29
    Calling the NYC AI research community 📣 Our third Research Meetup is almost here, and this time we're diving into benchmarking for coding agents. Join us at the @scale_AI NYC office to hear from researchers building at the frontier — plus food, drinks, and swag. Speakers:
  • user avatar
    Scale Labs
    @ScaleAILabs
    Jul 23
    Today we're releasing the complete EnigmaEval dataset, our benchmark built from sophisticated puzzle-hunt problems in partnership with @CAIS. We believe open access to EnigmaEval will help push reasoning and evaluation research forward. So far, Claude Fable 5 and GPT-5.6 Sol
    Image
  • user avatar
    Scale Labs
    @ScaleAILabs
    Jul 23
    Frontier-Bench (formerly Terminal-Bench) by @harborframework, an open-source benchmark that stress-tests AI agents on real command-line tasks, is now live. We're proud to be the top contributor to the launch set with tasks designed to pinpoint exactly where an agent breaks.
    user avatar
    Ryan Marten
    @ryan_marten
    Jul 23
    We’re releasing Frontier-Bench: a benchmark that measures and evolves with the frontier of agent work. Built by the team behind Terminal-Bench and Harbor, Frontier-Bench is an on-going community effort. Frontier-Bench v0.1 contains 74 tasks on which the best agents score ~34%
    Image

Log in or sign up for X

See what’s happening and join the conversation

Continue with phone
or
Log in with username or email
Terms·Privacy·Cookies·Accessibility·Ads Info·© 2026 X Corp.
Advertisement
Advertisement