JUST ADDED: @OpenAI GPT-6 Astra, @AnthropicAI Claude Fable 5.1, and @GoogleDeepMind Gemini 3.8 Flash just joined our leaderboards.
Check out the updated rankings:
An AI agent can top a benchmark and still be undeployable.
Introducing READY (Reliable Enterprise Agent Deployment), the first evaluation framework for qualifying AI agents for deployment on real enterprise workflows.
We’re inviting the research community to build it with us.
The response to RSI Bench has been incredible.
Hundreds of researchers and builders filled out our interest form, showing just how much energy there is around pushing the frontier of what it means for an agent to do research.
We’re incredibly excited to bring together some of
📣Call for contributions + co-authorship!
RSI Bench is our ongoing effort to evaluate whether AI agents can develop the capabilities required to advance AI R&D.
Got a frontier AI research problem on your mind? Submit it to RSI Bench, and set the standard the whole industry uses
Clinical AI agents can be right for the wrong reasons.
Real clinical work requires more than just medical knowledge. Agents need to navigate longitudinal records, reconcile evidence, ground conclusions, and know when to abstain.
We introduce CliniCARE-Bench, a new benchmark
Introducing HarnessOpt-Bench: A model improving the code of another AI agent is a concrete, measurable slice of recursive self-improvement (RSI).
So which frontier models are actually good at it? We built a head-to-head benchmark to find out. 🧵