New Harbor docs!
We've released many exciting features that were previously undocumented like simulating users, streaming, and regrading trials.
Check them out or give your coding agent the MCP.
docs.harborframework.com
We are releasing AutoResearchExam, a benchmark on open-ended machine learning and engineering tasks. Our benchmark covers seven research areas including model training, data curation, AI safety and interpretability.
benchmarks.bespokelabs.ai/autoresearchex…
In each task, we give agents 24 hours
Worked with Mercor Research and the SkyRL team, training Qwen3.5-397B-A17B on APEX-Agents (long-horizon office work) off-the-shelf data with SkyRL, improving Pass@1 from 16% to 27%.
The post is more of a practical field guide for what to do given an RL dataset, de-risking step
We've pushed a version update to the Terminal-Bench dataset and leaderboard.
Terminal-Bench 4.0 calibrates task resources (time, cpu, memory), implements task fixes, and removes saturated tasks.
Some agents need a sandbox where they can do work, liking editing files, installing dependencies, and launching builds. To score these agents you run the tests and check what changed, which takes a clean container per attempt.
Harbor is a Python framework for specifying