← Warmer Sun
Research dossier

Questions and Boundaries

Source-linked research behind the Altman and Epoch AI price comparison.

26 September 2026 · Warmer Sun

Primary questions

  1. What did Altman fix? A broad capability level, illustrated by GPT-4 and GPT-4o; the numerical example is token price, not job cost. [1]
  2. What did Epoch fix? A score threshold on each named benchmark, then the cheapest estimated model/settings bill to reach it on each date. [2]
  3. Is 13× a forecast? No. It is an annualized summary of about three years of observed/fitted cost-frontier declines. Continuation is uncertain. [2]
  4. Does the 725× GPQA pair set that rate? No. It is a single threshold and exceptionally steep interval; the reported average pools five benchmarks and many score levels. [2, 4]
  5. Why can the meters diverge? Reasoning tokens per question, performance of each token, model mix, and optimization over settings all change the cost of an answer beyond a token’s sticker price. [2]
  6. How robust are the observations? The 75% science-example rows are in Figure 2; Epoch cautions that model coverage and prices are noisy, and tasks can be optimized for benchmark questions. [2, 4]
  7. What would change the decision to keep separate tiles? A credible common dataset for both token price and full cost-to-score, with comparable score, models, date ranges, and uncertainty, could support a linked multi-meter chart—but not simple value splicing.

Unresolved questions

  • How much of Epoch’s measured decline transfers to longitudinal, quality-checked real work rather than benchmark questions? The report does not establish a general conversion.
  • What is the cost of switching models, integrating them, retrying failures, and reviewing answers for a particular buyer? The report explicitly focuses on a cheapest frontier, not a complete enterprise bill.
  • Does the same slope persist beyond the September 2026 observation window? Not established; changes in test difficulty, model supply, energy, or deployment costs could alter it.
  • How much of the decline is hardware efficiency versus algorithms, pricing strategy, competition, and more efficient reasoning? The report does not causally decompose the aggregate rate.

Source details