A long-form essay
The Reviewer Still Wins
And that is why AI pays — GDPval, the honest oversight math, and the 2026 scoreboard.
Last updated: 2026-08-12
In one breath. AI can draft real professional work. Serious companies will still have a person check it. That still saves money.
OpenAI’s GDPval test is a blind taste test for finished office work—not a trivia quiz. Experts pick which package they would rather ship: the human’s or the AI’s. The flashy headline says AI is about 100× faster and cheaper. That number only counts the computer. It ignores the checker.
Count a checker, and redo the whole job when the AI loses, and a strong model still wins on average: about 12% less time and 15% less cost than starting from a blank page. More if your reviews are faster. Below about a 28% “AI beats human” rate, mandatory review can make AI a net drag. Above it, “we still need people” is true—and “so skip AI” is not.
For decades, big technology has had the same awkward habit. In 1987 the economist Robert Solow joked that you could see computers everywhere except in the productivity numbers. Factories had screens. Offices had terminals. The national accounts barely moved.
Later, the historian Paul David compared that lag to electricity. Early factories often bolted electric motors onto layouts built for steam—one big engine, belts everywhere. The new power source sat on the old floor plan. Only when plants were redesigned around motors at each machine did productivity show up clearly.
That is the AI argument in knowledge work today. The demos are real. The “hours saved” spreadsheets often pretend the floor is already rebuilt. Someone still has to check the brief. Someone still signs. Someone still owns the bad number.
Then OpenAI published a test that put that awkwardness into one page of honest math.
They called it GDPval.
Not a quiz
GDPval is not an exam. It is a taste test for finished work.

OpenAI started from the biggest pieces of the U.S. economy, then picked high-pay jobs that are mostly digital knowledge work—lawyers, software developers, nurses, engineers, and peers across 44 occupations. Those roles sit on roughly three trillion dollars a year in wages in the authors’ framing.
Veterans with about fourteen years of experience wrote real assignments—with the messy attachments real work has. They also produced the gold answer: their own package. The full set has 1,320 tasks; 220 are public. A typical gold task is about seven hours of expert work. That length matters. Seven hours is long enough for structure, taste, formatting, and the quiet choices that separate a dump of text from something a partner might send.
The blindfold

Other experts from the same jobs get the request and two finished packages. They do not get a reliable label saying which is human and which is AI. They choose. A careful comparison takes more than an hour. That hour is the cost of judgment—the cost organizations erase when they only show stopwatch generation time.
Win rate means: how often the AI package beats the human gold.
In the 2025 study, the best overall model was judged as good as or better than the human nearly half the time. The stricter “AI wins” rate used in the money math was about 39% for a strong model. An older, weaker model won only 12.5% of the time. Strengths split: one model looked better; another was more accurate. Models still lose by ignoring instructions, skipping files, formatting badly, or inventing a number. Shorter tasks favor AI more than multi-day monsters. The average hides a map by job and sector.
Tools and checklists matter too. When researchers made models inspect their own slides and PDFs, quality jumped. A plain chat box under-samples what this test measures with tools turned on.
The hundredfold that is not a department
Here is the number that escaped into headlines: AI finishes these tasks roughly 100× faster and cheaper than experts.
![]()
OpenAI said the quiet part out loud. That number is pure computer time and API price. It does not include a human check, a fix, or putting the work into the real process. Silicon is cheap. Judgment is not.
If you stop at 100×, you get two bad religions:
- Fire the team.
- AI is a toy, because someone still checks.
Both fail the same test: what happens when a reviewer stays in the loop.
The math in plain English
You do not need advanced math. Here is the process real organizations can actually run:
- AI drafts (minutes, cheap).
- An expert reviews (in the study, first-time graders spent about two hours; your internal checklist can be faster).
- If the AI fails the quality bar, a human does the full job from scratch (about seven hours in the study).
Name the pieces once:
| Symbol | Meaning | Study average |
|---|---|---|
| HT | Human time to create the whole package | ~404 minutes (~7 hours) |
| RT | Human time to review an AI draft | ~109 minutes (~2 hours) |
| MT | Model time to draft | ~a few minutes |
| w | Win rate — share of times AI beats the human gold | 39% for a strong model |
| 1 − w | Lose rate — share of times you fully redo | 61% when w = 39% |
Same idea for dollars: HC, RC, MC are the matching labor/API cost proxies (~$361 human create, ~$86 review, under $1 for the model).
Always vs sometimes
- Always: draft + review → MT + RT
- Sometimes (only when AI loses): add full human create → (1−w)·HT
So expected time is just:
Read it left to right: model draft, plus review, plus lose-rate times full human time.
Expected cost is the same shape:
Speed vs doing it yourself is human time divided by that expected time:
That is why “someone still has to check” does not kill the case. You are not paying a full human job on every task—only on the fraction the AI loses.

Plug in the numbers (strong 2025 model)
Use w = 0.39, so 1−w = 0.61, with MT ≈ 4.5, RT = 109, HT = 404:
Same path on cost: about 1.18× cheaper (~$310 expected vs ~$361 alone). Roughly 45 minutes and $50 saved per average task under harsh rules.
| Path | Time | Labor-cost proxy |
|---|---|---|
| Human alone | ~7 hours | ~$360 |
| AI + review + full redo when AI loses | ~6 hours expected | ~$310 |
| Saved | ~45 minutes (~12%) | ~$50 (~15%) |
If the expert can try the model several times before giving up, the paper’s upper band is about 1.4× faster and 1.6× cheaper. Real life is often better still: people usually edit a weak draft instead of throwing it away. The study counted full rebuilds, so these savings are a lower bound.
The weaker model at 12.5% wins (1−w = 0.875) was a net slowdown under the same formula. Mandatory review of a bad model is not wise caution. It is a tax.
Break-even: when does review-still-required AI beat blank page?
AI-with-review is faster when expected time is less than human-alone time:
which rearranges to one clean rule:
In words: win rate must beat (draft time + review time) divided by full human create time.

With the study averages:
- Below ~28%: the shrug is right for this process.
- Above ~28%: the shrug leaves money on the table.
- Hold w = 39% and cut review from ~109 minutes to 30 minutes, and the same model’s speed-up jumps from about 1.12× toward about 1.4×.
You own how fast safe review is, and which tasks you send. The vendor owns how often the draft is good enough (w).
Nine people who were never fired
Picture one hundred experts. Each finishes one “seven-hour class” task per working day for a year.

- All human: on the order of 148,000 hours.
- AI + review + redo when needed: about 132,000 hours.
- Difference: roughly 16,000 hours—about nine full-time people of capacity.
- On the study’s wage proxies: on the order of $1 million+ a year.
- Every draft still reviewed.
Those nine people are not a firing plan. They are backlog cleared, more clients served, exceptions handled, juniors taught judgment instead of first-draft grind. Treat the savings only as headcount cuts and you miss the point twice: once by overfiring, once by underbuilding the review system that makes the math real.
The scoreboard kept moving
The 2025 paper is a snapshot. The same family of tasks kept getting harder for old models and easier for new ones.
By August 2026, a public leaderboard called GDPval-AA (run by Artificial Analysis) became a common race track. It uses a different scoring style than OpenAI’s human “win rate,” so you cannot paste a leaderboard number into the savings formula and call it done. As a direction signal, though, it is loud: top systems kept climbing, and Grok 4.6 jumped hard versus its prior version while staying relatively cheap per task on that harness. Computer cost is still pocket change next to expert time. The binding cost remains the reviewer’s clock.
Another yardstick still shows that even strong models fail many strict expert checklists. High leaderboard score is not “passes every rule every time.” That is why the reviewer stays.
Update the mental model without breaking it: the 2025 math teaches the mechanism; the 2026 boards teach which systems to bake off now—not last year’s chat demo.
What this test cannot prove
- Real work is a conversation. GDPval tasks are mostly one-shot.
- A task is not a job. High scores on packages do not empty whole floors.
- OpenAI built the yardstick and sells models. They also scored rivals and opened a gold set—still, trust your own bakeoff.
- A few bad errors can cost far more than average savings. High-stakes work needs tail controls, not only averages.
None of that brings back “review voids the business case.” It bounds the claim. The math is a lower-bound proof for strong models on this kind of work—not a universal ROI calculator for every workflow.
What to do
Stop bolting a chat window onto the old blank-page process and calling it strategy.
- List the recurring digital packages that look like this test.
- Run a small blind bakeoff with your own seniors—twenty or fifty tasks is enough. Include current strong models, not only last year’s default.
- Measure three numbers: how often AI wins, minutes for a safe review, how often you redo vs edit.
- Scale only the green task families, with a written checklist.
- Keep humans on ambiguity, stakes, and ugly exceptions.
The sticky line is short enough for a Post-it:
Nothing in AI workplace productivity claims makes sense except in the light of how often AI wins after a human check.
If your win rate sits below break-even for your review costs, you are funding a museum of demos. If it sits above, every month you keep starting from a blank page on those tasks, you are paying a tax to a superstition—the idea that a necessary reviewer makes the machine worthless.
The reviewer still wins.
That is why the machine pays.
References
Research appendix (this site)
- Research appendix home
- Research index
- What GDPval is
- Dataset & grading
- Quality results
- Productivity math
- Implications
- 2026 leaderboards & Grok 4.6
- Question ledger
- External primary sources
Primary technical sources
- OpenAI GDPval announcement
- GDPval paper (arXiv:2510.04374)
- GDPval-AA v2 leaderboard
- Grok 4.6 launch
Full dossier
Six research sections, the 2026 Grok 4.6 leaderboard refresh, sources, and the master question ledger — including the full Try-1× math.