Research · Abstract · map

Research index

Last updated: 2026-08-12

Executive abstract

GDPval is OpenAI’s 2025 evaluation of whether frontier AI models can produce real professional work products—briefs, plans, decks, spreadsheets, designs—across 44 knowledge occupations in the 9 largest U.S. GDP sectors. Experts with ~14 years’ average experience write the tasks; other experts blindly compare model vs human deliverables. On the public 220-task gold set, the best model (Claude Opus 4.1) was judged as good as or better than the human about 47.6% of the time; GPT-5’s stricter win rate in the cost tables is 39%. Raw generation looks like ~100× speed/cost versus humans, but that ignores oversight. When you model AI draft + expert review + full human redo on losses, GPT-5 is still about 1.12× faster and 1.18× cheaper than working alone (up to ~1.39× / 1.63× if you retry the model). Independently recomputed math matches the paper. Break-even win rate under gold-set averages is ~28%—below that (GPT-4o at 12.5%), mandatory review can make AI a net time sink; above it, paying people to review still pencils.

Through August 2026, third-party GDPval-AA v2 Elo scores (Human Baseline = 1000) show the same task family still climbing: Claude Opus 5 max ~1849, Grok 4.6 high ~1749–1753 (up from Grok 4.5 ~1526), GPT-5.6 Sol max ~1728. Those Elo numbers are not the paper’s human win rate w, but they are a leading signal that agentic quality kept rising—and that cost-efficient frontier options (Grok 4.6 at AA ~$0.84/task class pricing) widen the case for draft-first workflows if local win rate and review minutes still clear break-even. Confirmed: task-level peer pressure is real and still moving. Unverified for your firm: your local win rates, true internal review times, and multi-turn messy work. Implication: the question is no longer “must humans check AI?” (yes, in serious work) but “have we designed review and routing so the check is cheaper than blank-page creation—and are we baking off 2026-class models, not 2025 press releases?”

Short answers

Question Short answer
What is GDPval? Expert-graded benchmark of AI vs human deliverables on GDP-weighted knowledge jobs
Quality today? Best models near intermittent parity in 2025 paper; 2026 GDPval-AA Elo still climbing (Opus 5, Grok 4.6, Sol)
Is 100× real? Yes for inference only; no as an org throughput claim
With human review? GPT-5 still ~12% time / 15% cost expected savings (Try 1×); more if review is faster or w higher
Still hire reviewers? Yes—and that can still beat pure human on average once \(w\gtrsim 28\%\)
Grok 4.6? ~1750 GDPval-AA v2 Elo (#3 on AA table); not a substitute for local human w
Jobs over? No. Tasks ≠ jobs; ambiguity & tails remain human

Core mechanism

Economic weight selects occupations → experts encode real work as multi-file tasks → models produce deliverables → experts pick winners → win rate w plugs into expected-time formulas with review cost → productivity with oversight.

← Appendix hubWhat GDPval is →