We have been building the "virtual lab" mentioned in (anthropic.com/institute/recu…) and letting agents iterate on it. It will be important for measuring RSI progress and for accelerating automated safety research
We've set up @AISecurityInst's Inspect platform as a lightweight remote eval service: Postgres-backed job queue, Dockerized API + worker pool, remote submission, and the existing Inspect web UI for viewing runs. It keeps the core Inspect workflow intact, but makes it practical to
We're releasing 'ai-sft' on @huggingface: a 34GB supervised fine-tuning dataset for training models on AI research tasks. 2.7M examples spanning research code generation, scientific QA, and technical problem solving, built from our research-focused data collections. Each example
We're releasing 's2orc-safety' on @huggingface: a AI safety slice of our s2orc-enriched dataset with 16,806 papers across jailbreaks, prompt injection, red teaming, model security, privacy, robustness, alignment, and more.
Each paper is enriched with structured fields for