❓ Does Epoch confirm, extend, or contradict Altman’s observation?
It supports the direction, extends the measurement, and neither verifies nor refutes the exact original pair. Altman’s example asks how much one token costs from a model representing a comparable broad capability; Epoch asks how many dollars it costs to obtain a particular test score using the cheapest model and reasoning budget available. A token is a billed input or output unit; the score is an outcome. Holding a score fixed improves comparability for that one test, but does not make all kinds of intelligence interchangeable. [1, 2]
| Dimension | Altman, February 2025 | Epoch, September 2026 |
|---|---|---|
| Meter | Price per token in a named model comparison | Estimated total dollars per question to achieve at least a chosen test score |
| Target kept constant | Broad “given level” represented by a GPT-4/GPT-4o pair | Explicit threshold on each of five benchmarks |
| Rate | About 10× cheaper every 12 months (named observation) | About 47% cheaper quarterly since 2023, roughly 13× annually (modeled average) |
| Example | Altman’s about 150× token-price pair, early 2023–mid 2024 | Epoch’s roughly 725× GPQA pair, January 2025–July 2026 |
| Bias/limit | Pair and model choice; token price is not cost per completed task | Benchmark selection, cheapest-model assumption, reasoning-budget estimation, incomplete/noisy data |
A 47% quarterly decline means 53% of the previous quarter’s cost remains; compounding four quarters gives roughly 12.7× annual cost reduction. The 10× and 13× rates are the same order of magnitude, but an exact numerical match is neither expected nor required. The dramatic worked pairs do not define either long-run rate. [1, 2]
Mechanism: The cost of a task depends on the price of each token and how much work the model does to meet the bar. Since reasoning models can spend more tokens to improve accuracy, a falling token price can coexist with a higher bill for a hard answer. Conversely a better or more efficient model can sharply cut the bill even if its token sticker price barely changes. This is why Epoch explicitly contrasts earlier token-price studies with its cost-to-achieve-performance measure. [2]
Boundary: Epoch’s frontier is an optimistic shopping exercise, not an enterprise procurement basket. A buyer can face integration cost, output checks, latency, failures, and other constraints absent from a benchmark price. Benchmark gains can also outpace real-work improvements. Conversely, the study may miss cheap models or settings not evaluated. Its authors call the three-year data incomplete and noisy. Treat the slope as a dated observation and compare it against a repeatable workload before using it to forecast budget savings. [2]
Decision rule: Keep two plots with their own Y-axis units. Altman is a named law-of-the-shape in an indexed token-price story; Epoch is a research estimate of a cost-to-score frontier. Do not overwrite one with the other, and do not label the science-question result as the price of intelligence in general.