Harbor Hub
Find and share Harbor tasks for training and evaluation.
| Dataset | Tasks |
|---|---|
terminal-bench/terminal-bench Terminal-Bench is a benchmark for measuring agents' abilities to complete tasks using a terminal. | 66 |
terminal-bench-science/terminal-bench-science A benchmark for evaluating AI agents on research workflows across the life, physical, earth, mathematical, and engineering sciences. | 70 |
terminal-bench/terminal-bench-2-1 Version 2.1 of Terminal-Bench, a benchmark for evaluating agents in terminal environments. | 89 |
datacurve/deep-swe-1-1 DeepSWE: Measuring frontier coding agents on original, long-horizon engineering tasks | 113 |
abundant/swe-marathon Ultra long-horizon software engineering tasks from SWE-Marathon. | 20 |
scale-ai/swe-atlas-qna SWE-Atlas - Codebase QnA is a benchmark of deep codebase comprehension and QnA problems for coding agents. Checkout https://github.com/scaleapi/SWE-Atlas/ for instructions on running it. | 124 |
Publish your first dataset
Add the Harbor skill to your coding agent, then run /publish.
npx skills add harbor-framework/harbor --skill publish/publish