← Warmer Sun
Research dossier

What Epoch Measured

Source-linked research behind the Altman and Epoch AI price comparison.

26 September 2026 · Warmer Sun

❓ What is held fixed, which costs enter the bill, and how is the rate obtained?

Luke Emberson and David Roodman’s 22 September 2026 report asks for the cheapest available way to reach a given score on a given test at each date. Its five primary benchmarks are FrontierMath Tiers 1–3, OTIS Mock AIME mathematics questions, GPQA Diamond science questions, chess puzzles, and Mystery Game Puzzles. Those tests cover different sorts of performance, not every task a business might buy. The report’s underlying dataset draws on more models and evaluations than the four chart dots in the Future Forge draft. [2, 4]

The core quantity is a cost frontier: for a chosen score, look across eligible models and settings available by a date and take the lowest estimated cost of achieving at least that score. The price is dollars for the run or question rather than the price of one token. A model’s accuracy can change with its thinking effort and allowed token budget. Epoch follows a method from the Center for AI Standards and Innovation to estimate a model’s performance at different budget limits from a run transcript, and adds targeted evaluations of cheap and open-weight models missing from earlier benchmarking. For some open models it estimates rented-hardware costs, not an API invoice; a comparison on five models available both ways found cost differences of less than 30%. [2]

The report finds that the price of a given score fell about 47% per quarter since 2023, or about 13× per year, averaged across its five primary benchmarks. That is a fitted description of the measured period. On game puzzles it reports about 39–43% per quarter; on mathematics about 50–52%. For freshly state-of-the-art scores, its average is 66% per quarter; for a score reached two years earlier, 32% per quarter. The rate is thus sensitive to which score and when it first became attainable. [2]

Classification: measured, approximately exponential decline in a frontier price over about three years. The cost frontier itself changes in steps as models arrive and prices change; the annualized slope smooths those steps. A plausible mechanism is a combination of more efficient chips and models, different reasoning budgets, and competition for a once-novel capability. The report presents the premium on new scores and later competition as a possible explanation, not a proven causal decomposition. Constraints include benchmark gaming, switching costs, incomplete evaluations, latency and reliability. No evidence in the report guarantees a similar slope after 2026; a new paradigm might improve task-level efficiency, but its timing is not established. [2]