Research · Full question map
Question ledger
Last updated: 2026-08-12
Primary sources: OpenAI blog (2025-09-25); Patwardhan et al., arXiv:2510.04374 (2025-10-05)
Revision note
v1 built ledger-first for the user request: explain GDPval, do the productivity math, and show why human review of AI output can still leave large economic gains.
Primary questions (12)
P1. What is GDPval, and what problem was it built to solve?
- S1.1 How does it differ from MMLU, SWE-Bench, or Humanity’s Last Exam?
- S1.2 Why name it after GDP?
- S1.3 What does “economically valuable task” mean operationally?
- S1.4 Is the unit of analysis the task, the occupation, or the job?
- S1.5 Who built it and when was it released?
- S1.6 What is open-sourced vs held private?
- RH: How does Artificial Analysis GDPval-AA relate to the original eval?
P2. What is inside the dataset?
- S2.1 How many tasks, occupations, sectors?
- S2.2 How were occupations chosen (GDP share, wages, digital-task share)?
- S2.3 Who wrote the tasks (experience bar, employers)?
- S2.4 What do reference files and deliverable types look like?
- S2.5 How long does a typical task take a human expert?
- S2.6 What quality-control pipeline was used?
- RH: Does “predominantly digital / knowledge work (≥60% digital tasks)” bias results upward for AI?
P3. How is model performance graded?
- S3.1 What is pairwise expert comparison?
- S3.2 What is a win, tie, and loss?
- S3.3 How long does human grading take?
- S3.4 How good is the automated grader?
- S3.5 Can graders still detect model style (blinding limits)?
- S3.6 Why prefer win rate over a saturating score?
P4. What do the headline quality results say?
- S4.1 Which models were tested at launch?
- S4.2 What was Claude Opus 4.1’s win/tie rate?
- S4.3 What was GPT-5’s win rate used in the cost tables?
- S4.4 How fast did OpenAI models improve GPT-4o → GPT-5?
- S4.5 Where do models win vs lose by sector, occupation, file type, task duration?
- S4.6 What failure modes dominate (instruction-following, formatting, hallucination)?
- RH: Does aesthetic strength (Claude) inflate economic value vs accuracy strength (GPT-5)?
P5. What is the naive speed/cost claim (~100×), and why is it incomplete?
- S5.1 What does “naive” ratio measure (HT/MT, HC/MC)?
- S5.2 Why does OpenAI still publish it?
- S5.3 What workplace costs does it omit?
- S5.4 When would naive 100× be nearly true in practice?
P6. What is the human-oversight productivity math? (core)
- S6.1 Define HT, HC, RT, RC, MT, MC, w.
- S6.2 Derive Try-1× expected time and cost.
- S6.3 Derive Try-n× and the n→∞ limit.
- S6.4 Reproduce Table 2 numbers for GPT-5 and GPT-4o.
- S6.5 What is the break-even win rate under Try-1×?
- S6.6 How sensitive are gains to review-time RT?
- RH: Does the paper over-penalize models by freezing w across retries?
P7. Even if you must pay people to review AI output, when does AI still pay?
- S7.1 Walk a single average gold-set task in dollars and hours.
- S7.2 Scale to a 100-expert knowledge team for a year.
- S7.3 Compare “always human” vs “AI draft + review + redo if fail.”
- S7.4 When does low win rate make AI net negative (GPT-4o case)?
- S7.5 How does routing only high-win tasks change ROI?
- S7.6 What about catastrophic error tails not in the average?
P8. What organizational operating model does the math imply?
- S8.1 Draft-first vs decide-first workflows?
- S8.2 Where should human time move (judgment, exceptions, client ambiguity)?
- S8.3 What metrics should a PMO track (win rate by task family, RT, redo rate)?
- S8.4 How does scaffolding / reasoning effort change economics?
- S8.5 What changes for regulated industries (finance, health, government)?
P9. What are the strongest limitations and critiques?
- S9.1 One-shot vs multi-draft real work?
- S9.2 Well-scoped prompts vs ambiguous real requests?
- S9.3 Tasks ≠ jobs (occupation-level displacement fallacy)?
- S9.4 Self-reported human times and BLS median wages?
- S9.5 Vendor-authored eval and incentive bias?
- S9.6 Later third-party leaderboards (GDPval-AA) vs original paper protocol?
P10. How should the trend be classified?
- S10.1 Is quality improvement on GDPval exponential, linear, or stepwise?
- S10.2 What mechanism drives gains (scale, reasoning effort, scaffolding, multimodal tools)?
- S10.3 What bottlenecks remain (instruction-following, long-horizon agency, ambiguity)?
- S10.4 What would falsify “approaching expert parity”?
P11. What are the second-order economic implications?
- S11.1 Complement vs substitute for expert labor?
- S11.2 Wage and task-mix effects inside knowledge occupations?
- S11.3 Why GDP macro studies can lag capability evals?
- S11.4 What does “up elevator” language commit institutions to operationally?
P12. What should a decision-maker do with GDPval numbers this quarter?
- S12.1 Which internal task families map cleanly to GDPval-style work?
- S12.2 How to run a local mini-GDPval with own experts?
- S12.3 Act / watch / ignore thresholds on win rate and RT?
- S12.4 What not to conclude (full job replacement, universal 100× ROI)?
P13. How do 2026 live leaderboards (esp. Grok 4.6) change the picture?
- S13.1 What is GDPval-AA v2 vs OpenAI paper protocol?
- S13.2 Where does Grok 4.6 sit vs Opus 5, Fable 5, GPT-5.6 Sol?
- S13.3 How large was the Grok 4.5 → 4.6 jump?
- S13.4 Can AA Elo replace paper w in Try-1× math?
- S13.5 What does AA cost-per-task imply next to human HC?
- S13.6 What does Snorkel GDPval+ pass-rate evidence add as a caution?
Unresolved
| Question | Why open |
|---|---|
| Exact Claude / Gemini / Grok cost ratios under same oversight model (2025 paper) | Paper: cost estimates not available for non-OpenAI models at analysis time |
| True production RT for experienced internal reviewers (vs 109 min first-time graders) | Paper uses first-time grading RT; production may be faster |
| Occupation-level win rates as a public numeric table | Figures in paper; full numeric export not fully extracted here |
| Independent reproduction of gold-set pairwise grades | Automated grader ~66% agreement; full human regrade expensive |
| How much multi-turn client ambiguity would cut effective w | Explicit future-work item; not measured in v1 |
| Mapping AA Elo → OpenAI-style human win rate w | No published calibration; protocols differ |
| Independent human pairwise regrade of Grok 4.6 on gold set | Not public as of 2026-08-12 package update |
Stabilization note
Ledger v1.1 adds P13 after Aug 2026 Grok 4.6 / GDPval-AA refresh. Core 2025 math unchanged. Further rounds only if user wants occupation heatmaps or Elo→w calibration studies.