❓ Which actual cost-record observations support the science-example staircase, and do they define the 13× average?
Data transcription: Epoch’s public Figure 2 CSV, filtered to bench=gpqa, level=0.75. The date is the date field in the figure; cost is the CSV’s dollars-per-question estimate, not an audited invoice. The recorded accuracy is at least 75% in each row. These are all eight cost-record entries in that filtered figure, not an interpolated annual index. [4]
| Figure date | Model / setting | Estimated accuracy | Estimated cost per question (USD) |
|---|---|---|---|
| 2025-01-31 | o3 (high) | 75.76% | $0.298906 |
| 2025-03-25 | Gemini 2.5 Pro (Jun 2025) | 80.36% | $0.089116 |
| 2025-06-05 | Gemini 2.5 Pro (Jun 2025) | 77.77% | $0.012932 |
| 2025-12-01 | DeepSeek-V3.2 | 77.72% | $0.008821 |
| 2026-02-13 | Qwen3.5 397B-A17B | 79.8% | $0.007102 |
| 2026-02-25 | Qwen 3.5 Flash (hosted 35B-A3B) | 76.24% | $0.0021805 |
| 2026-06-01 | MiniMax-M3 | 75.08% | $0.001213 |
| 2026-07-09 | GPT-5.6 Luna (low) | 75.25% | $0.000412 |
The Gemini row’s parenthesized “Jun 2025” is the model label in the source CSV; its first date column is March 2025. The 75% target is a threshold, not an assertion that every model achieved exactly 75%. Later rows are more affordable ways of meeting or exceeding it under the report’s costing method.
The first and last observations support Epoch’s headline about 725× cheaper in under 18 months when compared at full source precision. The draft Future Forge GPQA tile (local ID ai-price-of-thought) picks four observations as milestone dots and anchors a separate modeled 13×-per-year line to the first point. It is not in the live catalog. The pairwise 725× fall is much steeper than the five-benchmark average and must not be presented as that average’s fitted line. [2, 4]