Research · Full question map

Question ledger

Last updated: 2026-08-12

Primary sources: OpenAI blog (2025-09-25); Patwardhan et al., arXiv:2510.04374 (2025-10-05)

Revision note

v1 built ledger-first for the user request: explain GDPval, do the productivity math, and show why human review of AI output can still leave large economic gains.


Primary questions (12)

P1. What is GDPval, and what problem was it built to solve?

P2. What is inside the dataset?

P3. How is model performance graded?

P4. What do the headline quality results say?

P5. What is the naive speed/cost claim (~100×), and why is it incomplete?

P6. What is the human-oversight productivity math? (core)

P7. Even if you must pay people to review AI output, when does AI still pay?

P8. What organizational operating model does the math imply?

P9. What are the strongest limitations and critiques?

P10. How should the trend be classified?

P11. What are the second-order economic implications?

P12. What should a decision-maker do with GDPval numbers this quarter?

P13. How do 2026 live leaderboards (esp. Grok 4.6) change the picture?


Unresolved

Question Why open
Exact Claude / Gemini / Grok cost ratios under same oversight model (2025 paper) Paper: cost estimates not available for non-OpenAI models at analysis time
True production RT for experienced internal reviewers (vs 109 min first-time graders) Paper uses first-time grading RT; production may be faster
Occupation-level win rates as a public numeric table Figures in paper; full numeric export not fully extracted here
Independent reproduction of gold-set pairwise grades Automated grader ~66% agreement; full human regrade expensive
How much multi-turn client ambiguity would cut effective w Explicit future-work item; not measured in v1
Mapping AA Elo → OpenAI-style human win rate w No published calibration; protocols differ
Independent human pairwise regrade of Grok 4.6 on gold set Not public as of 2026-08-12 package update

Stabilization note

Ledger v1.1 adds P13 after Aug 2026 Grok 4.6 / GDPval-AA refresh. Core 2025 math unchanged. Further rounds only if user wants occupation heatmaps or Elo→w calibration studies.

← Research referencesBack to essay →