JobBench: Aligning Agent Work with Human Will

Measuring agents by GDP alone asks how much of a human's job can be taken away.

JobBench asks how much of that job can be given back — built on the work that experts across real-world professions actually want delegated to AI.

agent_01
Current leader
Muse Spark 1.2
Meta
Weighted score61.6%
0
Professions
0
Tasks
0
Criteria

In collaboration with

University of Washington
UC Santa Barbara
Stanford University
Carnegie Mellon University
University of Notre Dame
IBM Research
BakeAI
Michigan State University
UC Berkeley
Northwestern University
University of Chicago

Adopted by

JobBench has been adopted by Meta's Muse Spark 1.1, Moonshot's Kimi K3, and Alibaba's Qwen 3.8 Max

ImageMuse Spark 1.1Meta
ImageKimi K3Moonshot
ImageQwen 3.8 MaxAlibaba
§ 01 — Why Human Will

Economics alone is not enough.

The conversation about AI in the workplace has been framed almost entirely in economic terms: what fraction of working hours can agents absorb? how much of GDP is exposed to automation? Benchmarks like OpenAI's GDPval inherit this framing by design — they select tasks that represent economic value, and score agents on whether they can deliver the professional knowledge output.

We believe this framing, on its own, is not enough.

If agents are going to share the professional workplace with humans, the question is not only what work is most economically valuable to automate, but what work do the humans in that role actually want automated? This is a humanist problem. It treats the professional not as labor to be displaced, but as a collaborator whose judgment about their own craft matters — and it is the premise JobBench is built on.

The economic question

GDPval

OpenAI

“What fraction of a human's job is economically valuable to automate?”

The humanist question

JobBench

Ours

“What work do the humans in that role actually want automated?”

Read the full essay
§ 02 — Rankings

Model leaderboard

Family
1
ImageMuse Spark 1.2
61.6
2
ImageClaude Fable 5
57.4
3
ImageMuse Spark 1.1
54.7
04
ImageKimi K3
54.3
05
ImageQwen 3.8 Max
52.7
06
ImageClaude Opus 4.8
48.4
07
ImageGPT-5.6 SOL
45.4
08
ImageClaude Opus 4.7
44.5
09
ImageGLM 5.2
43.4
10
ImageGPT-5.5
38.3
11
ImageClaude Sonnet 4.6
36.6
12
ImageGPT-5.4
32.2
13
ImageGemini 3.5 Flash
31.5
14
ImageGPT-5.2
26.6
15
ImageClaude Sonnet 4.5
20.7
16
ImageGemini 3.1 Pro
15.9

All models run on the same harness — OpenCode v1.14.18 — with corresponding max reasoning effort. Grok 4.3 is used as the rubric judge. (Earlier runs used Grok 4.1 Fast, since retired by xAI, so scores may differ slightly from previously reported results.) The Grok judges are chosen mainly for cost: one full eval pass costs ~$2 with Grok 4.1 Fast and ~$20 with Grok 4.3. Claude Fable 5 runs with fallback to Claude Opus 4.8 on refusals.

§ 03 — Headroom

Far from saturation

GPT-5.4 — Codex CLI
GDPvalsaturating
83.0
JobBench61 pts headroom
38.9
GPT-5.2 Codex
70.9/24.8
GPT-5.3 Codex
70.9/33.7
GPT-5.4 — Codex CLI
83.0/38.9
Workload
JobBench over GDPval
Wall-clock per task2.40×
Tool calls per task1.40×
Trajectory lines1.40×
§ 04 — Methodology

From knowledge delivery to professional reasoning

§ 05 — Inside a task

What the agent is actually up against

Every JobBench task is a small dossier. Pick one role to see the details.

Role

Reporter — Connecticut investigative desk

Automation desire
4.00/5
Lead in Connecticut drinking water. The state says zero water hazards. The FOIA data says otherwise.
6 sources· 4 types·3 contradictions
Source flow
Heterogeneous inputs

Multiple Hartford-area systems exceed the 15 ppb federal action level.

conflicts with CT_2024_Surveillance_ReportFOIA exceedances vs. 0% home-hazard finding
conflicts with EPA_LCRI_FactsheetRule finalized vs. current enforcement cycle

0% of investigated homes identified water as a lead hazard.

conflicts with FOIA_water_dataFOIA exceedances vs. 0% home-hazard finding

CT rows only for 2017–2019; 2020–2022 are dagger-marked non-submissions.

conflicts with martinez_interviewCDC n=1,666 vs. Martinez 30% clinic-specific

10 ppb action level finalized Oct 2024 — not yet enforceable.

conflicts with FOIA_water_dataRule finalized vs. current enforcement cycle

Pediatric referrals up 30% post-threshold change (Dr. Martinez).

conflicts with CDC_2017_2022_Blood_LeadCDC n=1,666 vs. Martinez 30% clinic-specific

Waterbury 16.1 ppb vs. Newark 47.9 ppb — trajectory, not point-in-time.

Agent
reasoning over reporter sources
Deliverables
  • Thesis-driven pitch memo
  • 3-sheet data workbook
  • 15+ entry source verification log

Reasoning challenges by design

click for full detail
§ 06 — Breakdown

Heatmap

Harness: OpenCode v1.14.18 for all models; judge: Grok 4.3. “–” = the run produced no output there — the agent hit the per-task time limit, or the model refused (Claude Fable 5 cells otherwise include its Opus 4.8 fallback on refusals).

Scale0–10%10–20%20–30%30–40%40%+
Occupation
Fable557.4
MuseSpark 1.154.7
KimiK354.3
Qwen3.8 Max52.7
Opus4.848.4
GPT-5.6SOL45.4
Opus4.744.5
GLM5.243.4
GPT-5.538.3
Sonnet4.636.6
Gemini3.5 Flash31.5
Sonnet4.520.7
Gemini3.1 Pro15.9
Business / Financial Ops
Bookkeeping & Accounting Clerks7751264653260243227617
HR Specialists699442550383875289900
Licensing Examiners / Inspectors677872727269538150672588
Management Analysts61393830291744293326936
Personal Financial Advisors31312179314667810183100
Purchasing Agents6561566349484735355431207
Training & Development Specialists4761575854544949622626317
Avg.5947455347464640353323107
Office / Admin Support
Court Clerks58553429372426421002918
Customer Service Reps66535066535063294229880
Data Entry Keyers89886763705766725028593849
Medical Secretaries5454545454464141315441158
Police / Fire Dispatchers68363636573619572847173628
Secretaries & Admin Assistants5239373038274831272353110
Avg.66555047544146423437382619
Computer / Mathematical
Biostatisticians2357513823374346442549520
CS Researchers354337342029241824179255
Statisticians43504453505340534955483123
User Support Specialists6553575955384861371539944
Web Administrators36724884366060723636242412
Avg.41554754374443503829341921
Architecture / Engineering
Civil Engineers61535863415049504445323131
Mechanical Eng. Technicians6147705045483936353029203
Mechanical Engineers55456742555836451803390
Petroleum Engineers525652285232322832321200
Avg.5750624648473940322727159
Management
Financial Managers67769073555239332850432924
Health Services Managers4539333938213220143416814
IT / IS Managers46544858285443274330251710
Supply Chain Managers29503517538518225555
Avg.47555147324230252730221413
Arts / Media
Producers1006478691008356787847531922
Reporters & Correspondents73576373736063634723331020
Technical Writers65666864594553624456435517
Avg.79627069776357685642432820
Other (Legal · Sales · Science · Edu.)
Lawyers75757563637563383850255013
Online Merchants6760777069705059506035619
Securities Sales Agents272741413835145127240140
Soc. Sci. Research Assistants74717570797064737074664642
Sociology Teachers (Postsec.)51635255604254353639401315
Tech & Sci. Sales Reps7339634633334435251419811
Avg.61566457575448494144312317

Cite

@misc{li2026jobbenchaligningagentwork,
  title         = {JobBench: Aligning Agent Work With Human Will},
  author        = {Yuetai Li and Yichen Feng and Zhangchen Xu and Zixian Ma and others},
  year          = {2026},
  eprint        = {2605.26329},
  archivePrefix = {arXiv},
  primaryClass  = {cs.AI},
  url           = {https://arxiv.org/abs/2605.26329}
}