Skip to main content

Better Harness · Open-source insights for the Agent Work Loop

Delegate coding to agents. Improve the loop around them.

Better Harness turns project and session evidence into loop-level insights, prioritized improvements, and verifiable next steps—inside the coding agent you already use.

  • Open source · MIT
  • Host-specific setup
  • Missing evidence stays explicit
Sample finding · evidence-boundedBetter Harness sample HTML report showing an evidence-bounded finding with its impact, expected output, scoped AI fix, and acceptance checksEvidence, impact, bounded repair, and acceptance checks in one reviewable report.

Choose your coding agent

Ten host adapters are supported. Six have verified setup paths; Pi, Kimi Code, WorkBuddy, and Grok link to their current support boundaries.

Follow delivery from intent to commit

Harness Inspector traces product intent through agent activity, sessions, files, and commits in one read-only workspace, keeping evidence strength and limitations visible.

Harness Inspector session view: a synchronized timeline of prompts, tool calls, and commits with the Evidence Drawer explaining each link

The interactive sample uses fictional English data. It does not read your workspace, Git history, or coding-agent sessions.

Open the interactive sample Read the architecture

Turn evidence into the next concrete improvement

Better Harness keeps unsupported claims out of the score and turns observed workflow gaps into findings a team can inspect, discuss, and verify.

Visible evidence

See which project or session signal supports each finding.

Prioritized impact

Start with the workflow gap that matters most.

Bounded repair

Keep the proposed change scoped to the observed problem.

Acceptance checks

Know what evidence would make the improvement reviewable.

Explore the self-contained English sample report

Track recorded change over time

Static final frame of Better Harness report history showing five Agent Work Loop dimensions over time

This static final frame summarizes historical Harness reports. It shows recorded trends, not causal proof of improvement.

How Better Harness works

Better Harness combines feedforward guides (AGENTS.md, specs, Skills, acceptance criteria) with feedback sensors (linters, tests, Hooks, evaluation agents), and evaluates five parts of delivery—the Agent Work Loop:

Task Understanding

Does the agent know the goal and what “done” means?

Controlled Execution

Is the work on supported, repeatable paths?

Change Validation

Is there evidence the change actually works?

Reliable Delivery

Does AI speed bypass quality checks or acceptance?

Learning Capture

Does the next task benefit from this one?

Better Harness architecture: six public Quickstart hosts plus Pi, Kimi Code, WorkBuddy, and Grok adapter support feed three independent evidence agents, unified analysis, host-neutral outputs, and repair

Ten capability-level host adapters feed the same evidence pipeline. Six have verified Quickstart paths; Pi, Kimi Code, WorkBuddy, and Grok keep their current adapter-support boundaries explicit.