Research · Abstract · map
Research index
Last updated: 2026-08-12
Executive abstract
GDPval is OpenAI’s 2025 evaluation of whether frontier AI models can produce real professional work products—briefs, plans, decks, spreadsheets, designs—across 44 knowledge occupations in the 9 largest U.S. GDP sectors. Experts with ~14 years’ average experience write the tasks; other experts blindly compare model vs human deliverables. On the public 220-task gold set, the best model (Claude Opus 4.1) was judged as good as or better than the human about 47.6% of the time; GPT-5’s stricter win rate in the cost tables is 39%. Raw generation looks like ~100× speed/cost versus humans, but that ignores oversight. When you model AI draft + expert review + full human redo on losses, GPT-5 is still about 1.12× faster and 1.18× cheaper than working alone (up to ~1.39× / 1.63× if you retry the model). Independently recomputed math matches the paper. Break-even win rate under gold-set averages is ~28%—below that (GPT-4o at 12.5%), mandatory review can make AI a net time sink; above it, paying people to review still pencils.
Through August 2026, third-party GDPval-AA v2 Elo scores (Human Baseline = 1000) show the same task family still climbing: Claude Opus 5 max ~1849, Grok 4.6 high ~1749–1753 (up from Grok 4.5 ~1526), GPT-5.6 Sol max ~1728. Those Elo numbers are not the paper’s human win rate w, but they are a leading signal that agentic quality kept rising—and that cost-efficient frontier options (Grok 4.6 at AA ~$0.84/task class pricing) widen the case for draft-first workflows if local win rate and review minutes still clear break-even. Confirmed: task-level peer pressure is real and still moving. Unverified for your firm: your local win rates, true internal review times, and multi-turn messy work. Implication: the question is no longer “must humans check AI?” (yes, in serious work) but “have we designed review and routing so the check is cheaper than blank-page creation—and are we baking off 2026-class models, not 2025 press releases?”
Short answers
| Question | Short answer |
|---|---|
| What is GDPval? | Expert-graded benchmark of AI vs human deliverables on GDP-weighted knowledge jobs |
| Quality today? | Best models near intermittent parity in 2025 paper; 2026 GDPval-AA Elo still climbing (Opus 5, Grok 4.6, Sol) |
| Is 100× real? | Yes for inference only; no as an org throughput claim |
| With human review? | GPT-5 still ~12% time / 15% cost expected savings (Try 1×); more if review is faster or w higher |
| Still hire reviewers? | Yes—and that can still beat pure human on average once \(w\gtrsim 28\%\) |
| Grok 4.6? | ~1750 GDPval-AA v2 Elo (#3 on AA table); not a substitute for local human w |
| Jobs over? | No. Tasks ≠ jobs; ambiguity & tails remain human |
Core mechanism
Economic weight selects occupations → experts encode real work as multi-file tasks → models produce deliverables → experts pick winners → win rate w plugs into expected-time formulas with review cost → productivity with oversight.