Research · Near expert parity

Quality results

Last updated: 2026-08-12

Quality results: are models near experts?

On real professional deliverables, how close are frontier models to industry experts—and where do they still fail?

Headline (gold set, human pairwise grades)

Models evaluated in the launch paper include GPT-4o, o4-mini, o3, GPT-5, Claude Opus 4.1, Gemini 2.5 Pro, and Grok 4.

Result Figure Source framing
Claude Opus 4.1 better or as good as human 47.6% Best overall in set; strong on aesthetics (layout, slides, formatting)
Models matching/beating humans “Just over half the tasks” in the strongest reading of better-or-equal Approaching parity, not past it as a default
GPT-5 win rate used in cost table 39.0% Stricter “better than human” win rate w
GPT-4o win rate 12.5% Far from parity; economics collapse under oversight
Trajectory GPT-4o → GPT-5 Roughly linear rise on gold set over ~1 year; blog: performance more than doubled / more than tripled depending on metric wording Clear capability progress, not a one-model fluke

Split of strengths: GPT-5 relatively stronger on accuracy (instruction care, calculations, domain facts). Claude Opus 4.1 relatively stronger on aesthetics and many non-plain-text file types (PDF, slides, spreadsheets). Pure text tasks show lower win rates than polished multi-file packages—important when someone claims “AI writes better than lawyers” from a slide-deck score.

Where models lose

Expert justification clustering (paper Fig. 8 theme):

Shorter tasks (0–2 hours human) show higher model win rates than multi-day monsters. Sector and occupation heatmaps vary: some areas near parity, others still cold. That heterogeneity is the operational message: do not average your company into one ROI number.

Scaffolding and reasoning move the needle

Controlled extras in the paper:

So GDPval is not only a model IQ test. It is also a test of harness quality—tools, file pipelines, and checklist prompts. Enterprises that deploy “chat box only” will under-sample the capability the paper measures with tools enabled.

Trend classification (Trend-analysis Rule)

Trend: Frontier model pairwise quality vs expert humans on GDPval gold tasks
Period: roughly 2024 (GPT-4o) → 2025 (GPT-5 / Opus 4.1) in the paper’s OpenAI series plot
Pattern: Authors describe improvement as roughly linear in time on this metric—not a measured multi-decade exponential with a stable doubling time.
Mechanism: combination of stronger base models, more test-time reasoning, multimodal file tools, and scaffolding—not a single knob.
Bottlenecks: instruction-following, long-horizon coherence, formatting reliability, ambiguity handling (out of scope for v1), catastrophic tail errors.
Next paradigm: interactive multi-turn GDPval (explicitly planned) may re-open the gap or widen model advantage depending on whether models use feedback better than one-shot.
Reach: Even a linear climb on expert-relative win rate is economically huge because each point of w feeds the oversight productivity formulas in 04.

Classification label: too short a window to call exponential; best current label is rapid, roughly linear climb on a preference metric, with discrete jumps when scaffolding/tools change. Do not smuggle LOAR language onto a one-year win-rate chart without longer series.

What “approaching expert parity” does not mean

It does not mean:

It does mean: on a large, expert-built sample of real digital deliverables, the best models are no longer embarrassing amateurs. They are intermittent peers—sometimes preferred, often close, still frequently worse—especially when nobody reviews them.

← Dataset & gradingProductivity math →