Research · Near expert parity

Quality results

Last updated: 2026-08-12

Quality results: are models near experts?

❓ On real professional deliverables, how close are frontier models to industry experts—and where do they still fail?

Headline (gold set, human pairwise grades)

Models evaluated in the launch paper include GPT-4o, o4-mini, o3, GPT-5, Claude Opus 4.1, Gemini 2.5 Pro, and Grok 4.

Result Figure Source framing
Claude Opus 4.1 better or as good as human 47.6% Best overall in set; strong on aesthetics (layout, slides, formatting)
Models matching/beating humans “Just over half the tasks” in the strongest reading of better-or-equal Approaching parity, not past it as a default
GPT-5 win rate used in cost table 39.0% Stricter “better than human” win rate w
GPT-4o win rate 12.5% Far from parity; economics collapse under oversight
Trajectory GPT-4o → GPT-5 Roughly linear rise on gold set over ~1 year; blog: performance more than doubled / more than tripled depending on metric wording Clear capability progress, not a one-model fluke

Split of strengths: GPT-5 relatively stronger on accuracy (instruction care, calculations, domain facts). Claude Opus 4.1 relatively stronger on aesthetics and many non-plain-text file types (PDF, slides, spreadsheets). Pure text tasks show lower win rates than polished multi-file packages—important when someone claims “AI writes better than lawyers” from a slide-deck score.

Where models lose

Expert justification clustering (paper Fig. 8 theme):

Shorter tasks (0–2 hours human) show higher model win rates than multi-day monsters. Sector and occupation heatmaps vary: some areas near parity, others still cold. That heterogeneity is the operational message: do not average your company into one ROI number.

Scaffolding and reasoning move the needle

Controlled extras in the paper:

So GDPval is not only a model IQ test. It is also a test of harness quality—tools, file pipelines, and checklist prompts. Enterprises that deploy “chat box only” will under-sample the capability the paper measures with tools enabled.

Trend classification (Trend-analysis Rule)

Trend: Frontier model pairwise quality vs expert humans on GDPval gold tasks
Period: roughly 2024 (GPT-4o) → 2025 (GPT-5 / Opus 4.1) in the paper’s OpenAI series plot
Pattern: Authors describe improvement as roughly linear in time on this metric—not a measured multi-decade exponential with a stable doubling time.
Mechanism: combination of stronger base models, more test-time reasoning, multimodal file tools, and scaffolding—not a single knob.
Bottlenecks: instruction-following, long-horizon coherence, formatting reliability, ambiguity handling (out of scope for v1), catastrophic tail errors.
Next paradigm: interactive multi-turn GDPval (explicitly planned) may re-open the gap or widen model advantage depending on whether models use feedback better than one-shot.
Reach: Even a linear climb on expert-relative win rate is economically huge because each point of w feeds the oversight productivity formulas in 04.

Classification label: too short a window to call exponential; best current label is rapid, roughly linear climb on a preference metric, with discrete jumps when scaffolding/tools change. Do not smuggle LOAR language onto a one-year win-rate chart without longer series.

What “approaching expert parity” does not mean

It does not mean:

It does mean: on a large, expert-built sample of real digital deliverables, the best models are no longer embarrassing amateurs. They are intermittent peers—sometimes preferred, often close, still frequently worse—especially when nobody reviews them.

← Dataset & gradingProductivity math →