JobBench: Aligning Agent Work with Human Will

Measuring agents by GDP alone asks how much of a human's job can be taken away.

JobBench asks how much of that job can be given back — built on the work that experts across real-world professions actually want delegated to AI.

agent_01
Current leader
Claude Opus 5
Anthropic
Weighted score65.7%
0
Professions
0
Tasks
0
Criteria

In collaboration with

University of Washington
UC Santa Barbara
Stanford University
Carnegie Mellon University
University of Notre Dame
IBM Research
BakeAI
Michigan State University
UC Berkeley
Northwestern University
University of Chicago

Adopted by

JobBench has been adopted by Meta's Muse Spark 1.1, Moonshot's Kimi K3, and Alibaba's Qwen 3.8 Max

ImageMuse Spark 1.1Meta
ImageKimi K3Moonshot
ImageQwen 3.8 MaxAlibaba
§ 01 — Why Human Will

Economics alone is not enough.

The conversation about AI in the workplace has been framed almost entirely in economic terms: what fraction of working hours can agents absorb? how much of GDP is exposed to automation? Benchmarks like OpenAI's GDPval inherit this framing by design — they select tasks that represent economic value, and score agents on whether they can deliver the professional knowledge output.

We believe this framing, on its own, is not enough.

If agents are going to share the professional workplace with humans, the question is not only what work is most economically valuable to automate, but what work do the humans in that role actually want automated? This is a humanist problem. It treats the professional not as labor to be displaced, but as a collaborator whose judgment about their own craft matters — and it is the premise JobBench is built on.

The economic question

GDPval

OpenAI

“What fraction of a human's job is economically valuable to automate?”

The humanist question

JobBench

Ours

“What work do the humans in that role actually want automated?”

Read the full essay
§ 02 — Rankings

Model leaderboard

Family
1
ImageClaude Opus 5
65.7
2
ImageMuse Spark 1.3
64.9
3
ImageMuse Spark 1.2
61.6
04
ImageGLM-5.3
61.4
05
ImageClaude Fable 5
57.4
06
ImageMuse Spark 1.1
54.7
07
ImageKimi K3
54.3
08
ImageQwen 3.8 Max
52.7
09
ImageClaude Opus 4.8
48.4
10
ImageGPT-5.6 SOL
45.4
11
ImageClaude Opus 4.7
44.5
12
ImageGLM 5.2
43.4
13
ImageGPT-5.5
38.3
14
ImageClaude Sonnet 4.6
36.6
15
ImageGPT-5.4
32.2
16
ImageGemini 3.5 Flash
31.5
17
ImageInkling
31.1
18
ImageNemotron 3 Ultra
27.2
19
ImageGPT-5.2
26.6
20
ImageClaude Sonnet 4.5
20.7
21
ImageGemini 3.1 Pro
15.9

All models run on the same harness — OpenCode v1.14.18 — with corresponding max reasoning effort. Grok 4.3 is used as the rubric judge. (Earlier runs used Grok 4.1 Fast, since retired by xAI, so scores may differ slightly from previously reported results.) The Grok judges are chosen mainly for cost: one full eval pass costs ~$2 with Grok 4.1 Fast and ~$20 with Grok 4.3. Claude Fable 5 runs with fallback to Claude Opus 4.8 on refusals.

§ 03 — Headroom

Far from saturation

GPT-5.4 — Codex CLI
GDPvalsaturating
83.0
JobBench61 pts headroom
38.9
GPT-5.2 Codex
70.9/24.8
GPT-5.3 Codex
70.9/33.7
GPT-5.4 — Codex CLI
83.0/38.9
Workload
JobBench over GDPval
Wall-clock per task2.40×
Tool calls per task1.40×
Trajectory lines1.40×
§ 04 — Methodology

From knowledge delivery to professional reasoning

§ 05 — Inside a task

What the agent is actually up against

Every JobBench task is a small dossier. Pick one role to see the details.

Role

Reporter — Connecticut investigative desk

Automation desire
4.00/5
Lead in Connecticut drinking water. The state says zero water hazards. The FOIA data says otherwise.
6 sources· 4 types·3 contradictions
Source flow
Heterogeneous inputs

Multiple Hartford-area systems exceed the 15 ppb federal action level.

conflicts with CT_2024_Surveillance_Report — FOIA exceedances vs. 0% home-hazard finding
conflicts with EPA_LCRI_Factsheet — Rule finalized vs. current enforcement cycle

0% of investigated homes identified water as a lead hazard.

conflicts with FOIA_water_data — FOIA exceedances vs. 0% home-hazard finding

CT rows only for 2017–2019; 2020–2022 are dagger-marked non-submissions.

conflicts with martinez_interview — CDC n=1,666 vs. Martinez 30% clinic-specific

10 ppb action level finalized Oct 2024 — not yet enforceable.

conflicts with FOIA_water_data — Rule finalized vs. current enforcement cycle

Pediatric referrals up 30% post-threshold change (Dr. Martinez).

conflicts with CDC_2017_2022_Blood_Lead — CDC n=1,666 vs. Martinez 30% clinic-specific

Waterbury 16.1 ppb vs. Newark 47.9 ppb — trajectory, not point-in-time.

Agent
reasoning over reporter sources
Deliverables
  • Thesis-driven pitch memo
  • 3-sheet data workbook
  • 15+ entry source verification log

Reasoning challenges by design

click for full detail
§ 06 — Breakdown

Heatmap

Harness: OpenCode v1.14.18 for all models; judge: Grok 4.3. “–” = the run produced no output there — the agent hit the per-task time limit, or the model refused (Claude Fable 5 cells otherwise include its Opus 4.8 fallback on refusals).

Scale0–10%10–20%20–30%30–40%40%+
Occupation
GLM-5.361.4
Fable557.4
MuseSpark 1.154.7
KimiK354.3
Qwen3.8 Max52.7
Opus4.848.4
GPT-5.6SOL45.4
Opus4.744.5
GLM5.243.4
GPT-5.538.3
Sonnet4.636.6
Gemini3.5 Flash31.5
Inkling31.1
Nemotron3 Ultra27.2
Sonnet4.520.7
Gemini3.1 Pro15.9
Business / Financial Ops
Bookkeeping & Accounting Clerks4977512646–532602432271317617
HR Specialists84699442550383875289956900
Licensing Examiners / Inspectors786778727272695381506725676788
Management Analysts34613938302917442933269181336
Personal Financial Advisors5431312179314667810183118800
Purchasing Agents6965615663494847353554313331207
Training & Development Specialists7147615758545449496226263050317
Avg.6359474553474646403533233428107
Office / Admin Support
Court Clerks45–585534–29372426421001602918
Customer Service Reps58665350665350632942298372980
Data Entry Keyers62898867637057667250285953533849
Medical Secretaries5454545454544641413154413326158
Police / Fire Dispatchers36683636365736195728471736473628
Secretaries & Admin Assistants4852393730382748312723537283110
Avg.50665550475441464234373835302619
Computer / Mathematical
Biostatisticians5823575138233743464425494242520
CS Researchers53354337342029241824179197255
Statisticians64435044535053405349554812193123
User Support Specialists5965535759553848613715394128944
Web Administrators60367248843660607236362424362412
Avg.59415547543744435038293427261921
Architecture / Engineering
Civil Engineers76615358634150495044453257483131
Mechanical Eng. Technicians7061477050454839363530292642203
Mechanical Engineers8255456742555836451803318090
Petroleum Engineers52525652285232322832321201200
Avg.7057506246484739403227272526159
Management
Financial Managers92677690735552393328504337142924
Health Services Managers5145393339382132201434163314814
IT / IS Managers39465448582854432743302512241710
Supply Chain Managers2329503517538518225523655
Avg.51475551473242302527302226141413
Arts / Media
Producers10010064786910083567878475347331922
Reporters & Correspondents73735763737360636347233333301020
Technical Writers64656668645945536244564334455517
Avg.79796270697763576856424338362820
Other (Legal · Sales · Science · Edu.)
Lawyers88757575636375633838502550505013
Online Merchants7067607770697050595060353223619
Securities Sales Agents542727414138351451272401411140
Soc. Sci. Research Assistants71747175707970647370746630334642
Sociology Teachers (Postsec.)62516352556042543536394051261315
Tech & Sci. Sales Reps6673396346333344352514191412811
Avg.68615664575754484941443132262317

Cite

@misc{li2026jobbenchaligningagentwork,
  title         = {JobBench: Aligning Agent Work With Human Will},
  author        = {Yuetai Li and Yichen Feng and Zhangchen Xu and Zixian Ma and others},
  year          = {2026},
  eprint        = {2605.26329},
  archivePrefix = {arXiv},
  primaryClass  = {cs.AI},
  url           = {https://arxiv.org/abs/2605.26329}
}