Skip to content

eval(headless): DeepSeek V4 Flash Terminal-Bench 2.1 harness A/B — Maka vs OpenCode #1680

Description

@Astro-Han

Goal

Produce reproducible evidence of whether Maka or OpenCode is more effective and resource-efficient when both use DeepSeek V4 Flash under the same Terminal-Bench 2.1 conditions.

This is a same-model harness comparison, not a model bake-off.

Frozen experiment design

  • Benchmark: Terminal-Bench 2.1, all 89 tasks
  • Task revision: d49e28f1e4ddd13d289e85a5f312a66750951932
  • Task-tree fingerprint: sha256:456826aa4c47ed309716c964c96d2a3acc998764ebc84f3e8449c807d74bd4e7
  • Arms: Maka vs OpenCode 1.17.18
  • Model: deepseek/deepseek-v4-flash
  • Reasoning effort: max
  • Attempt policy: paired Pass@1
  • Execution policy: up to four task pairs concurrently, with Maka and OpenCode started in parallel inside each pair; at most eight cells may run concurrently
  • Both arms use the same model, provider credential, task order, task instruction, deadline, verifier, and retry policy while retaining their native harness behavior.
  • Retry only infrastructure-invalid cells. Agent failures remain benchmark outcomes.
  • Credentials stay behind the host-side provider proxy and must not enter task containers or published artifacts.
  • The experiment configuration is frozen from the first accepted full-v5 result onward.

Current run

  • Run ID: deepseek-v4-flash-maka-vs-opencode-tbench-2.1-full-v5
  • Started: 2026-07-31
  • Status: full 89-pair run in progress
  • Pair concurrency: 4
  • Maximum concurrent cells: 8
  • Arm execution: parallel
  • The immutable manifest and append-only run journal have been created.
  • Runs v1–v4 are preserved as superseded or infrastructure-invalid evidence and are excluded from the accepted v5 comparison.
  • No external Oracle registry snapshot is configured. Oracle evidence is advisory; the official Terminal-Bench verifier remains the scoring authority.

Execution plan

  1. Run all 89 task pairs under the frozen v5 manifest.
  2. Continuously audit verifier validity, execution identity, provider usage, tool settlement, and infrastructure failures.
  3. Preserve invalid and superseded evidence; rerun only cells demonstrated to be infrastructure-invalid.
  4. Validate that paired results, usage accounting, JSON, CSV, and Markdown artifacts agree.
  5. Publish a report under docs/eval/.

Reported outcomes

Effectiveness:

  • Pass count and pass rate by arm
  • Paired wins, losses, and ties
  • Paired uncertainty and significance
  • Deadline and budget-exhaustion outcomes
  • Failure classification for discordant pairs

Resource efficiency:

  • Uncached input, cached input, output, and total tokens
  • Cache hit rate where available
  • API-equivalent cost using the pricing identity captured by the manifest
  • Wall-clock duration
  • Tokens, cost, and duration per pass

Completion criteria

  • Exactly 89 accepted task pairs are evaluated.
  • Every accepted cell has valid execution identity and verifier evidence.
  • JSON, CSV, and Markdown artifacts agree.
  • The final report records the frozen experiment identity and known limitations.
  • No credential appears in git, logs, commands, screenshots, fixtures, or published artifacts.

Related

中文翻译

目标

在相同的 Terminal-Bench 2.1 条件下,让 Maka 与 OpenCode 使用同一个 DeepSeek V4 Flash 模型,产出可复现的效果与资源效率对比证据。

这是同模型、不同 harness 的比较,不是模型横评。

冻结的实验设计

  • 基准:Terminal-Bench 2.1,共 89 个任务
  • 任务版本:d49e28f1e4ddd13d289e85a5f312a66750951932
  • 任务树指纹:sha256:456826aa4c47ed309716c964c96d2a3acc998764ebc84f3e8449c807d74bd4e7
  • 两个实验臂:Maka 与 OpenCode 1.17.18
  • 模型:deepseek/deepseek-v4-flash
  • 推理强度:max
  • 尝试策略:配对 Pass@1
  • 执行策略:最多同时运行 4 对任务;每一对内部同时启动 Maka 与 OpenCode,最多并发运行 8 个 cell
  • 两个实验臂使用相同的模型、提供方凭证、任务顺序、任务指令、时限、验证器与重试策略,同时保留各自原生的 harness 行为。
  • 只重试基础设施无效的 cell;Agent 自身失败仍计为基准结果。
  • 凭证仅保留在宿主机侧的 provider proxy 后方,不得进入任务容器或发布产物。
  • full-v5 的第一个有效结果开始冻结实验配置。

当前运行

  • Run ID:deepseek-v4-flash-maka-vs-opencode-tbench-2.1-full-v5
  • 启动日期:2026-07-31
  • 状态:完整 89 对任务正在运行
  • 配对并行度:4
  • 最大 cell 并行数:8
  • 两臂执行方式:并行
  • 不可变 manifest 与追加式运行日志已经创建。
  • v1–v4 均保留为已取代或基础设施无效的历史证据,不计入 v5 的有效比较。
  • 当前未配置外部 Oracle registry 快照。Oracle 证据仅作辅助,正式得分仍以 Terminal-Bench 官方 verifier 为准。

执行计划

  1. 按冻结的 v5 manifest 运行全部 89 对任务。
  2. 持续审计 verifier 有效性、执行身份、provider usage、工具收敛情况和基础设施故障。
  3. 保留无效和已被替代的原始证据;只重跑能够证明为基础设施无效的 cell。
  4. 验证配对结果、usage 计费以及 JSON、CSV、Markdown 产物相互一致。
  5. docs/eval/ 下发布报告。

报告指标

效果:

  • 各实验臂通过数与通过率
  • 配对胜、负、平
  • 配对不确定性与显著性
  • 超时与预算耗尽结果
  • 对结果不一致任务的失败机制分类

资源效率:

  • 非缓存输入、缓存输入、输出与总 token
  • 在数据可用时报告缓存命中率
  • 按 manifest 中冻结的计价身份计算 API 等价成本
  • 墙钟时间
  • 每次通过所需的 token、成本与时间

完成标准

  • 恰好完成 89 个有效任务对。
  • 每个有效 cell 都具备正确的执行身份和 verifier 证据。
  • JSON、CSV 与 Markdown 产物相互一致。
  • 最终报告记录冻结的实验身份与已知限制。
  • Git、日志、命令、截图、fixture 和发布产物中均不得出现凭证。

关联项

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions