← Warmer Sun
Research dossier

GPQA Diamond: The Cost Records

Source-linked research behind the Altman and Epoch AI price comparison.

26 September 2026 · Warmer Sun

❓ Which actual cost-record observations support the science-example staircase, and do they define the 13× average?

Data transcription: Epoch’s public Figure 2 CSV, filtered to bench=gpqa, level=0.75. The date is the date field in the figure; cost is the CSV’s dollars-per-question estimate, not an audited invoice. The recorded accuracy is at least 75% in each row. These are all eight cost-record entries in that filtered figure, not an interpolated annual index. [4]

Figure date Model / setting Estimated accuracy Estimated cost per question (USD)
2025-01-31 o3 (high) 75.76% $0.298906
2025-03-25 Gemini 2.5 Pro (Jun 2025) 80.36% $0.089116
2025-06-05 Gemini 2.5 Pro (Jun 2025) 77.77% $0.012932
2025-12-01 DeepSeek-V3.2 77.72% $0.008821
2026-02-13 Qwen3.5 397B-A17B 79.8% $0.007102
2026-02-25 Qwen 3.5 Flash (hosted 35B-A3B) 76.24% $0.0021805
2026-06-01 MiniMax-M3 75.08% $0.001213
2026-07-09 GPT-5.6 Luna (low) 75.25% $0.000412

The Gemini row’s parenthesized “Jun 2025” is the model label in the source CSV; its first date column is March 2025. The 75% target is a threshold, not an assertion that every model achieved exactly 75%. Later rows are more affordable ways of meeting or exceeding it under the report’s costing method.

The first and last observations support Epoch’s headline about 725× cheaper in under 18 months when compared at full source precision. The draft Future Forge GPQA tile (local ID ai-price-of-thought) picks four observations as milestone dots and anchors a separate modeled 13×-per-year line to the first point. It is not in the live catalog. The pairwise 725× fall is much steeper than the five-benchmark average and must not be presented as that average’s fitted line. [2, 4]