Harbor Hub

Find and share Harbor tasks for training and evaluation.

DatasetTasks
terminal-bench/terminal-bench

Terminal-Bench is a benchmark for measuring agents' abilities to complete tasks using a terminal.

66
terminal-bench-science/terminal-bench-science

A benchmark for evaluating AI agents on research workflows across the life, physical, earth, mathematical, and engineering sciences.

70
terminal-bench/terminal-bench-2-1

Version 2.1 of Terminal-Bench, a benchmark for evaluating agents in terminal environments.

89
datacurve/deep-swe-1-1

DeepSWE: Measuring frontier coding agents on original, long-horizon engineering tasks

113
abundant/swe-marathon

Ultra long-horizon software engineering tasks from SWE-Marathon.

20
scale-ai/swe-atlas-qna

SWE-Atlas - Codebase QnA is a benchmark of deep codebase comprehension and QnA problems for coding agents. Checkout https://github.com/scaleapi/SWE-Atlas/ for instructions on running it.

124

Publish your first dataset

Add the Harbor skill to your coding agent, then run /publish.

1. Install the Harbor skill
npx skills add harbor-framework/harbor --skill publish
2. Run /publish in your agent
/publish