← Warmer Sun
Essay · AI economics

Two Prices for the Same Intelligence

Altman priced tokens. Epoch priced the bill to reach a score. Both fall fast—but they are not the same meter.

26 September 2026 · Warmer Sun

Two prices for the same intelligence

In February 2025, Sam Altman put a startling number in a short post about the economics of AI. The cost to use a given level of intelligence, he wrote, was falling about tenfold every year. He pointed to the price of a token—the small piece of text an AI reads or writes—for GPT-4 in early 2023 and GPT-4o in mid-2024. On his comparison, it had fallen about 150-fold. [1]

A year and a half later, Luke Emberson and David Roodman at Epoch AI put a different receipt on the table. In January 2025, OpenAI’s o3 could score about 75 percent on GPQA Diamond, a demanding multiple-choice test in physics, chemistry, and biology. Epoch estimated that one question at that level cost about 30 cents. By July 2026, GPT-5.6 Luna could reach the same score for roughly four hundredths of a cent. About 725 times less. [2]

Both stories say that yesterday’s AI capability is getting cheaper. They do not charge for the same thing.

Illustration of the two meters used to price a given level of AI performance

The token meter

Altman’s observation fixes a broad level of intelligence and compares the price of access. His example uses a token, a piece of text consumed or produced by a model, as the billing unit. It is a memorable observation, not an audited price index: one pair of models, one executive’s comparison, and a claim of about 10× cheaper each year. A cheaper token is good news when the job uses about the same number of tokens.

But a token is an ingredient, not a finished answer. A model may need to write much more before it solves a hard problem. A restaurant could halve the price of flour without halving the price of dinner if every new recipe uses far more flour.

That is the gap Epoch measured.

The answer meter

Epoch held the test result fixed and asked which available model could achieve it for the lowest total bill. Its study covers five tests of mathematics, hard science, and games of skill. It estimates the cost of reaching the same score at different dates, including the model’s work to produce the answer. Across those tests, the cheapest route to a fixed level of performance fell about 47 percent per quarter from 2023 through 2026, equivalent to roughly 13× cheaper per year. This is a measured average over a short, changing set of tests—not a guarantee of next year’s price. [2]

The GPQA example makes the distinction tangible. Thirty cents and roughly $0.0004 buy the same test score, not the same amount of text. The 725-fold change is one especially steep pair over less than 18 months. It is not Epoch’s overall 13×-per-year rate, and it cannot simply be pasted onto Altman’s token-price chart. [2]

Why does the meter matter more now? A reasoning model can spend extra text-like steps working through a question. Such a model may be cheap per token but expensive per solved question. Another model may charge more per token yet get there in fewer steps. Epoch describes this distinction explicitly: prior studies often priced tokens from models capable of a score; its study estimates the bill to achieve that score. [2]

A rhyme, not a replication

The rates rhyme: Altman’s named observation says about 10× per year; Epoch’s estimated score-achievement frontier says about 13×. The second strengthens the broad claim that established AI abilities become dramatically cheaper. It also extends the claim into the reasoning era, when paying for a fixed amount of text is no longer a reliable stand-in for paying for a result.

It does not independently verify Altman’s 150-fold GPT-4-to-GPT-4o token comparison. Nor does a 13× average refute a 10× observation. Different periods, model sets, methods, and units can yield different numbers even when the underlying direction agrees. Epoch’s own results vary by test: game puzzles cheapened more slowly than mathematics. The price at a freshly reached score also tended to fall faster at first and slow later. [2]

There are limits to what this receipt buys. Epoch follows the cost frontier—the cheapest option available for a specified score—not what a typical customer actually selects. Its authors warn that benchmark scores may improve faster than useful work, and that their model coverage and prices are imperfect. A fixed score on a science quiz does not tell you what it costs to have a reliable scientist, nor what it costs to review a mistaken answer. [2]

If you are budgeting AI, ask for both prices. How much does the model charge per token? And how much does it cost, including all its work and retries, to achieve the result you need at a quality bar you can check? Altman named the falling price of access. Epoch measured the falling price of reaching a score.

The meter changes the story.


References

  1. Sam Altman, “Three Observations” (9 February 2025). Second observation and GPT-4/GPT-4o token example.
  2. Luke Emberson and David Roodman, “The plunging price of thought,” Epoch AI (22 September 2026). Main report, data, methods, examples, and limitations. Announcement thread.