Goal
Produce reproducible evidence of whether Maka or OpenCode is more effective and resource-efficient when both use DeepSeek V4 Flash under the same Terminal-Bench 2.1 conditions.
This is a same-model harness comparison, not a model bake-off.
Frozen experiment design
- Benchmark: Terminal-Bench 2.1, all 89 tasks
- Task revision:
d49e28f1e4ddd13d289e85a5f312a66750951932
- Task-tree fingerprint:
sha256:456826aa4c47ed309716c964c96d2a3acc998764ebc84f3e8449c807d74bd4e7
- Arms: Maka vs OpenCode
1.17.18
- Model:
deepseek/deepseek-v4-flash
- Reasoning effort:
max
- Attempt policy: paired Pass@1
- Execution policy: up to four task pairs concurrently, with Maka and OpenCode started in parallel inside each pair; at most eight cells may run concurrently
- Both arms use the same model, provider credential, task order, task instruction, deadline, verifier, and retry policy while retaining their native harness behavior.
- Retry only infrastructure-invalid cells. Agent failures remain benchmark outcomes.
- Credentials stay behind the host-side provider proxy and must not enter task containers or published artifacts.
- The experiment configuration is frozen from the first accepted
full-v5 result onward.
Current run
- Run ID:
deepseek-v4-flash-maka-vs-opencode-tbench-2.1-full-v5
- Started: 2026-07-31
- Status: full 89-pair run in progress
- Pair concurrency: 4
- Maximum concurrent cells: 8
- Arm execution: parallel
- The immutable manifest and append-only run journal have been created.
- Runs v1–v4 are preserved as superseded or infrastructure-invalid evidence and are excluded from the accepted v5 comparison.
- No external Oracle registry snapshot is configured. Oracle evidence is advisory; the official Terminal-Bench verifier remains the scoring authority.
Execution plan
- Run all 89 task pairs under the frozen v5 manifest.
- Continuously audit verifier validity, execution identity, provider usage, tool settlement, and infrastructure failures.
- Preserve invalid and superseded evidence; rerun only cells demonstrated to be infrastructure-invalid.
- Validate that paired results, usage accounting, JSON, CSV, and Markdown artifacts agree.
- Publish a report under
docs/eval/.
Reported outcomes
Effectiveness:
- Pass count and pass rate by arm
- Paired wins, losses, and ties
- Paired uncertainty and significance
- Deadline and budget-exhaustion outcomes
- Failure classification for discordant pairs
Resource efficiency:
- Uncached input, cached input, output, and total tokens
- Cache hit rate where available
- API-equivalent cost using the pricing identity captured by the manifest
- Wall-clock duration
- Tokens, cost, and duration per pass
Completion criteria
- Exactly 89 accepted task pairs are evaluated.
- Every accepted cell has valid execution identity and verifier evidence.
- JSON, CSV, and Markdown artifacts agree.
- The final report records the frozen experiment identity and known limitations.
- No credential appears in git, logs, commands, screenshots, fixtures, or published artifacts.
Related
中文翻译
目标
在相同的 Terminal-Bench 2.1 条件下,让 Maka 与 OpenCode 使用同一个 DeepSeek V4 Flash 模型,产出可复现的效果与资源效率对比证据。
这是同模型、不同 harness 的比较,不是模型横评。
冻结的实验设计
- 基准:Terminal-Bench 2.1,共 89 个任务
- 任务版本:
d49e28f1e4ddd13d289e85a5f312a66750951932
- 任务树指纹:
sha256:456826aa4c47ed309716c964c96d2a3acc998764ebc84f3e8449c807d74bd4e7
- 两个实验臂:Maka 与 OpenCode
1.17.18
- 模型:
deepseek/deepseek-v4-flash
- 推理强度:
max
- 尝试策略:配对 Pass@1
- 执行策略:最多同时运行 4 对任务;每一对内部同时启动 Maka 与 OpenCode,最多并发运行 8 个 cell
- 两个实验臂使用相同的模型、提供方凭证、任务顺序、任务指令、时限、验证器与重试策略,同时保留各自原生的 harness 行为。
- 只重试基础设施无效的 cell;Agent 自身失败仍计为基准结果。
- 凭证仅保留在宿主机侧的 provider proxy 后方,不得进入任务容器或发布产物。
- 从
full-v5 的第一个有效结果开始冻结实验配置。
当前运行
- Run ID:
deepseek-v4-flash-maka-vs-opencode-tbench-2.1-full-v5
- 启动日期:2026-07-31
- 状态:完整 89 对任务正在运行
- 配对并行度:4
- 最大 cell 并行数:8
- 两臂执行方式:并行
- 不可变 manifest 与追加式运行日志已经创建。
- v1–v4 均保留为已取代或基础设施无效的历史证据,不计入 v5 的有效比较。
- 当前未配置外部 Oracle registry 快照。Oracle 证据仅作辅助,正式得分仍以 Terminal-Bench 官方 verifier 为准。
执行计划
- 按冻结的 v5 manifest 运行全部 89 对任务。
- 持续审计 verifier 有效性、执行身份、provider usage、工具收敛情况和基础设施故障。
- 保留无效和已被替代的原始证据;只重跑能够证明为基础设施无效的 cell。
- 验证配对结果、usage 计费以及 JSON、CSV、Markdown 产物相互一致。
- 在
docs/eval/ 下发布报告。
报告指标
效果:
- 各实验臂通过数与通过率
- 配对胜、负、平
- 配对不确定性与显著性
- 超时与预算耗尽结果
- 对结果不一致任务的失败机制分类
资源效率:
- 非缓存输入、缓存输入、输出与总 token
- 在数据可用时报告缓存命中率
- 按 manifest 中冻结的计价身份计算 API 等价成本
- 墙钟时间
- 每次通过所需的 token、成本与时间
完成标准
- 恰好完成 89 个有效任务对。
- 每个有效 cell 都具备正确的执行身份和 verifier 证据。
- JSON、CSV 与 Markdown 产物相互一致。
- 最终报告记录冻结的实验身份与已知限制。
- Git、日志、命令、截图、fixture 和发布产物中均不得出现凭证。
关联项
Goal
Produce reproducible evidence of whether Maka or OpenCode is more effective and resource-efficient when both use DeepSeek V4 Flash under the same Terminal-Bench 2.1 conditions.
This is a same-model harness comparison, not a model bake-off.
Frozen experiment design
d49e28f1e4ddd13d289e85a5f312a66750951932sha256:456826aa4c47ed309716c964c96d2a3acc998764ebc84f3e8449c807d74bd4e71.17.18deepseek/deepseek-v4-flashmaxfull-v5result onward.Current run
deepseek-v4-flash-maka-vs-opencode-tbench-2.1-full-v5Execution plan
docs/eval/.Reported outcomes
Effectiveness:
Resource efficiency:
Completion criteria
Related
中文翻译
目标
在相同的 Terminal-Bench 2.1 条件下,让 Maka 与 OpenCode 使用同一个 DeepSeek V4 Flash 模型,产出可复现的效果与资源效率对比证据。
这是同模型、不同 harness 的比较,不是模型横评。
冻结的实验设计
d49e28f1e4ddd13d289e85a5f312a66750951932sha256:456826aa4c47ed309716c964c96d2a3acc998764ebc84f3e8449c807d74bd4e71.17.18deepseek/deepseek-v4-flashmaxfull-v5的第一个有效结果开始冻结实验配置。当前运行
deepseek-v4-flash-maka-vs-opencode-tbench-2.1-full-v5执行计划
docs/eval/下发布报告。报告指标
效果:
资源效率:
完成标准
关联项