Research · Near expert parity
Quality results
Last updated: 2026-08-12
Quality results: are models near experts?
❓ On real professional deliverables, how close are frontier models to industry experts—and where do they still fail?
Headline (gold set, human pairwise grades)
Models evaluated in the launch paper include GPT-4o, o4-mini, o3, GPT-5, Claude Opus 4.1, Gemini 2.5 Pro, and Grok 4.
| Result | Figure | Source framing |
|---|---|---|
| Claude Opus 4.1 better or as good as human | 47.6% | Best overall in set; strong on aesthetics (layout, slides, formatting) |
| Models matching/beating humans | “Just over half the tasks” in the strongest reading of better-or-equal | Approaching parity, not past it as a default |
| GPT-5 win rate used in cost table | 39.0% | Stricter “better than human” win rate w |
| GPT-4o win rate | 12.5% | Far from parity; economics collapse under oversight |
| Trajectory GPT-4o → GPT-5 | Roughly linear rise on gold set over ~1 year; blog: performance more than doubled / more than tripled depending on metric wording | Clear capability progress, not a one-model fluke |
Split of strengths: GPT-5 relatively stronger on accuracy (instruction care, calculations, domain facts). Claude Opus 4.1 relatively stronger on aesthetics and many non-plain-text file types (PDF, slides, spreadsheets). Pure text tasks show lower win rates than polished multi-file packages—important when someone claims “AI writes better than lawyers” from a slide-deck score.
Where models lose
Expert justification clustering (paper Fig. 8 theme):
- Instruction-following failures dominate for Claude / Gemini / Grok losses (wrong format, missing promised files, ignoring references).
- Formatting errors dominate GPT-5 losses more than pure instruction misses.
- All models still hallucinate or miscalculate sometimes.
- Failure severity is not binary: many losses are “acceptable but subpar”; a minority are bad/catastrophic (order-of-magnitude: secondary writeups cite ~29% of failure ratings in bad/catastrophic bands for GPT-5 analyses—treat as paper-appendix signal, use primary figures when auditing).
Shorter tasks (0–2 hours human) show higher model win rates than multi-day monsters. Sector and occupation heatmaps vary: some areas near parity, others still cold. That heterogeneity is the operational message: do not average your company into one ROI number.
Scaffolding and reasoning move the needle
Controlled extras in the paper:
- Higher reasoning effort improves o3 / GPT-5 GDPval scores.
- Better prompts + agent scaffolding (e.g., force multimodal self-inspection of PDFs/slides, best-of-N with a judge) raised win rates (~+5 percentage points in the prompt-tuning experiment) and cut egregious artifacts (black-square PDFs eliminated in their test; PowerPoint egregious formatting errors fell substantially).
So GDPval is not only a model IQ test. It is also a test of harness quality—tools, file pipelines, and checklist prompts. Enterprises that deploy “chat box only” will under-sample the capability the paper measures with tools enabled.
Trend classification (Trend-analysis Rule)
Trend: Frontier model pairwise quality vs expert humans on GDPval gold tasks
Period: roughly 2024 (GPT-4o) → 2025 (GPT-5 / Opus 4.1) in the paper’s OpenAI series plot
Pattern: Authors describe improvement as roughly linear in time on this metric—not a measured multi-decade exponential with a stable doubling time.
Mechanism: combination of stronger base models, more test-time reasoning, multimodal file tools, and scaffolding—not a single knob.
Bottlenecks: instruction-following, long-horizon coherence, formatting reliability, ambiguity handling (out of scope for v1), catastrophic tail errors.
Next paradigm: interactive multi-turn GDPval (explicitly planned) may re-open the gap or widen model advantage depending on whether models use feedback better than one-shot.
Reach: Even a linear climb on expert-relative win rate is economically huge because each point of w feeds the oversight productivity formulas in 04.
Classification label: too short a window to call exponential; best current label is rapid, roughly linear climb on a preference metric, with discrete jumps when scaffolding/tools change. Do not smuggle LOAR language onto a one-year win-rate chart without longer series.
What “approaching expert parity” does not mean
It does not mean:
- models already do entire jobs end-to-end unsupervised;
- every occupation is equally exposed;
- aesthetic wins equal fiduciary-grade accuracy;
- one-shot lab tasks equal messy organizational work.
It does mean: on a large, expert-built sample of real digital deliverables, the best models are no longer embarrassing amateurs. They are intermittent peers—sometimes preferred, often close, still frequently worse—especially when nobody reviews them.