Scale Labs

Research to Advance AI

Scale Labs advances AI through research. Our research focuses on agents, post-training, reasoning, safety, evaluation, and alignment, and the science of data.

[PAPERS]

Research papers and publications covering agents, post-training, reasoning, safety, evaluation, and alignment, and the science of data.

Date Title
9/4/2026
READY or Not: Reliable Enterprise Agent DeploymentAgents, Enterprise
8/27/2026
Spine-Branch Coordination for Multi-agent Computer UseAgents
8/20/2026
CliniCARE-Bench: Clinical Calibrated Audit of Medical Reasoning in EHREvaluation and Alignment
8/6/2026
HarnessOpt-Bench: Evaluating LLMs at Harness OptimizationAgents, Enterprise, Evaluation and Alignment
6/30/2026
DrugDiscoveryBench: Can Coding Agents Assist Early-Stage Drug Discovery?Agents, Enterprise, Evaluation and Alignment
6/29/2026
SWE-INTERACT: Reimagining SWE Benchmarks as User-Driven Long-Horizon Coding SessionsAgents, Evaluation and Alignment
View more

[BLOG]

Insights, analysis, and updates from Scale Labs

Evaluation and AlignmentSep 10, 2026

RUBRIC DROPOUT: A SIMPLE WAY TO MITIGATE REWARD HACKING IN RUBRIC-AS-REWARD RL

Rubric-based RL can quietly learn to game its reward: the training judge keeps assigning higher scores even as true quality declines. A one-line fix, inspired by neural-network dropout, mitigates the problem at virtually no additional cost.

AgentsSep 3, 2026

Introducing READY: What It Takes to Deploy an AI Agent

Today we're introducing READY (Reliable Enterprise Agent Deployment), a suite of industry-specific benchmarks that measure agents the way enterprises actually use them: working alongside people, inside real workflows.

AgentsSep 2, 2026

Who Grades the Graders? Rethinking Verifier Design for Computer Use Agents

As Computer Use Agents (CUA) take on complex professional tasks like processing emails, creating spreadsheets, drawing 3D diagrams, and drafting financial memos, verifiers are at the heart of providing evaluation and training signals around their capabilities. Benchmarks like OSWorld 2.0 and Agents' Last Exam use these verifiers, in the form of programmatic checks, to measure whether an agent is capable of completing production-grade work.

AgentsAug 26, 2026

CliniCARE-Bench: Clinical AI Agents Can Be Right for the Wrong Reasons

We evaluated 16 agentic systems on 25 clinical care scenarios over 750 real patient cases. Every single system committed to an answer more often than the evidence allowed, and up to one in five correct verdicts rested on an investigation the case authors had explicitly prohibited.

View allAll posts