feat(headless): add resumable GLM harness benchmark - #862
Merged
Conversation
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
Why
We need reproducible evidence for comparing harness effectiveness and economy without mixing model, prompt, task-order, or billing changes into the result. This extends the existing headless, fixed-prompt, and Harbor seams so a 40-task pilot can resume into the full 89-task suite without rerunning completed cells.
Closes #861.
Scope
This PR adds benchmark infrastructure only. It does not run the benchmark, commit result artifacts, add Kimi or Pi arms, introduce LiteLLM, or change product defaults.
Verification
npm test --workspace @maka/headless— 822 tests passed after rebasing onto the latestorigin/main.npm run check:stale— passed.git diff --check origin/main...HEAD— passed.Impact
Headless benchmark operators gain a single fixed entrypoint with append-only resume state, execution attestation, and normalized reports. Existing product behavior and public defaults are unchanged; no migration is required.
Reviewer notes
Please focus on fairness invariants, resume correctness, secret isolation, OpenCode token accounting, and whether the implementation reuses the closest existing headless/Harbor seams without parallel state.
Ready for review
Verificationexplains why it is not applicable