Back
Business 10 min read - 18 Sept. 26 - Félix Monnier

Should we fear the explosion in token prices? Five years of data on the cost of generative AI

September 2026 - prices recorded on 3 September 2026.
When a company wishes to automate tasks using generative artificial intelligence (AI), it creates a dependency on a Large Language Model (LLM) provider and is exposed to the variation in the cost of tokens (the billing unit for these models) which it does not control. Many therefore hesitate when launching a project, for fear that this price will soar due to providers increasing their rates once customers are captive, or subsidised introductory prices that would rise when it's time to make a profit. Five years of data allow both to be tested.

The finding: at equal performance, the price is divided by ten each year, down to a floor

Comparing prices over time requires reasoning at equal quality. Standardised test suites measure a model's quality: MMLU for general knowledge, GPQA for PhD-level questions, SWE-Bench for code, and LMArena Elo rankings for human preference.
These measurements remain imperfect, but they are sufficient to define quality tiers and to ask the right question, that of the price of the cheapest model that achieves a given tier at a given time. The answer is unambiguous, and three independent studies measure it with different methods. Public tariffs extend their series until 2026.
Quality TierTrajectoryFactorFloor ReachedSource
MMLU 42 (2021)60 $ → 0,06 $ / M tokens in 3 years÷ 1 0000,06 $, Nov. 2024a16z, « LLMflation »
MMLU 65 (2022)20 $ → 0,07 $ in 23 months÷ 2800,07 $, Oct. 2024Stanford AI Index 2025
MMLU 86 (2023)37,50 $ → 0,11 $ in 37 months÷ 3300,11 $, Apr. 2026Epoch AI, public tariffs
MMLU 42 QualityMMLU 65 QualityMMLU 86 Quality
plancher : le matériel et l'électricité0,01 $0,1 $1 $10 $100 $20222023202420252026$ par million de tokens×10Qualité MMLU 420,06 $Qualité MMLU 650,07 $Qualité MMLU 860,11 $GPT-3 davinci · OpenAInov. 2021 · 60 $ / M tokensGPT-3 davinci, après baisse · OpenAIsept. 2022 · 20 $ / M tokensLlama 3.2 3B · Together AInov. 2024 · 0,06 $ / M tokensplusieurs modèles au plancher · hébergeurs open-weightmars 2026 · 0,06 $ / M tokenstext-davinci-003 · OpenAInov. 2022 · 20 $ / M tokensGPT-3.5 Turbo · OpenAImars 2023 · 1,63 $ / M tokensGPT-3.5 Turbo 1106 · OpenAInov. 2023 · 0,75 $ / M tokensGemini 1.5 Flash-8B · Googleoct. 2024 · 0,07 $ / M tokensGPT-4 (8k) · OpenAImars 2023 · 37,5 $ / M tokensGPT-4 Turbo · OpenAInov. 2023 · 15 $ / M tokensGPT-4o · OpenAImai 2024 · 7,5 $ / M tokensGemini 2.0 Flash · Googlefévr. 2025 · 0,18 $ / M tokensGemini 2.5 Flash-Lite · Googlejuil. 2025 · 0,17 $ / M tokensDeepSeek V4 Flash · DeepInfraavr. 2026 · 0,11 $ / M tokens
On a logarithmic scale. The three curves plunge in parallel for three years until reaching a plateau, where the cost of a sufficiently small model comes down to its hardware and electricity.
Epoch AI measures a median decline of a factor of fifty per year across six benchmarks. The conservative order of magnitude to remember is a tenfold division each year at constant quality, and no exception has appeared since 2021. This slope holds for approximately three years per tier, the time it takes to go from tens of dollars to a few cents per million tokens, then it stops. The GPT-3 tier is worth six cents in November 2024 and still six cents in March 2026, while the GPT-4 tier takes fourteen months to drop from eighteen to eleven cents.
This floor makes the future clear for anyone designing an application today. A sufficiently small model eventually costs only the price of its hardware and electricity, and the curve stabilises there. The targeted quality tier therefore loses a factor of one hundred to one thousand in three years, then it remains at a few cents per million tokens as long as the computational cost holds. Observers tracking the series now announce a rate of three to five per year until 2027 on the higher tiers, as the lower tiers complete their descent.
A numerical example illustrates what this represents. A document extraction pipeline consuming 30 million input tokens and 3 million output tokens each month cost approximately 1 080 dollars per month on GPT-4 in 2023. An equivalent quality model brings the same bill down to a little over 3 dollars per month in 2026, a cost divided by 330 for the same service. Cache discounting, which removes 90% of the price of repeated input, and batch processing discounting, which removes half, accumulate on top of this.
0 $300 $600 $900 $1 200 $2023202420252026$ par mois, à périmètre constant1 080 $/mois3 $/moisplancher tenu depuis juil. 2025GPT-4 (8k) · OpenAImars 2023 · 30 $ en entrée, 60 $ en sortieGPT-4 Turbo · OpenAInov. 2023 · 10 $ en entrée, 30 $ en sortieGPT-4o · OpenAImai 2024 · 5 $ en entrée, 15 $ en sortieGPT-4o, après reprix · OpenAIoct. 2024 · 2,5 $ en entrée, 10 $ en sortieGPT-4.1 mini · OpenAIavr. 2025 · 0,4 $ en entrée, 1,6 $ en sortieGemini 2.5 Flash-Lite · Googlejuil. 2025 · 0,1 $ en entrée, 0,4 $ en sortieDeepSeek V4 Flash · DeepInfraavr. 2026 · 0,09 $ en entrée, 0,18 $ en sortie
The same trajectory appears here on a linear scale, where the previous graph showed a rate and this one illustrates a bill. It applies the actual input and output prices of the cheapest GPT-4 quality model for the same monthly volume at each date, and it flattens out from summer 2025 because the bill has reached its floor.

Why no one can reverse this curve

A trend becomes reassuring once we understand why it holds, and three mechanisms lock this in.
  1. The open source floor belongs to everyone. At each quality tier, an open-weight model, such as Llama, Mistral, DeepSeek or Qwen, arrives twelve to twenty-one months after the proprietary pioneer, and this delay is shortening. Competition among dozens of cloud hosts sets its price, with AWS, Azure, GCP or Groq, whose rates descend to the cost of hardware and electricity. A proprietary provider that increased its prices would be out of the market. January 2025 illustrated this brutally, when DeepSeek offered a quality comparable to OpenAI's best reasoning model for approximately 96% less. The gap has become structural, with a median mixed price of fifty-three cents per million tokens for open-weight models in 2026, compared to three dollars for proprietary ones.
  2. Changing models costs almost nothing. OpenAI's API format has become a de facto standard, exposed identically by almost the entire market. In an application where model calls live behind an abstraction layer, migration amounts to changing a configuration and putting the replacement through an evaluation set. This substitutability disciplines prices better than negotiation.
  3. Current prices already generate a margin. According to SemiAnalysis, Anthropic's gross margin on its API exceeds 80%, and what loses money for large labs remains the unlimited public subscription, unlike the token charged per usage. No hidden losses are therefore waiting to be covered by a future increase. The market is even moving in the opposite direction, as Google lowered the price of Gemini Flash by 78% in August 2024, in direct reaction to an OpenAI launch. OpenAI did the same on 30 July 2026, cutting the price of GPT-5.6 Luna by 80%, reducing it from one dollar to twenty cents per million input tokens.
Two physical drivers are added to these market mechanisms. Algorithmic efficiency doubles performance per parameter approximately every 3.3 months, and hardware follows the same trajectory, with an H100 GPU falling from approximately eight dollars per hour at the end of 2024 to less than two dollars among the cheapest hosts in 2026.

The four nuances to know

The absolute frontier, however, becomes more expensive. The cost of the best model of the moment increases by a factor of three to eighteen per year, according to the MIT study « The Price of Progress », because each marginal gain requires more computation. September 2026 prices show this, with GPT-6 Astra at ten dollars per million input tokens and fifty for output, while Pro tiers rise to thirty and one hundred eighty dollars. Maintaining the quality level of one's application follows a downward trajectory, whereas chasing the latest model can follow an upward trajectory.
The number of tokens per task increases. Reasoning models charge for invisible intermediate tokens, and agentic architectures consume ten to one hundred times more than a simple call. With a constant scope, the cost per task decreases, and it rises again as soon as the scope expands.
Energy and hardware slow the slope without straightening it. The strain on HBM memory lasts until 2027, electricity costs more, and hyperscalers are investing around 700 billion dollars in 2026. These frictions are real, and the floor holds despite them, supported by ever smaller models on ever more efficient hardware.
The real risk is withdrawal from the catalogue. Providers leave the prices of their older models in place and withdraw these models from the catalogue, with observed notice periods of approximately six months at OpenAI and at least sixty days at Anthropic. The example is ongoing, as OpenAI announced on 22 April 2026 the grouped withdrawal of GPT-3.5 Turbo, GPT-4, GPT-4 Turbo, o1, o3-mini and o4-mini by 23 October 2026. For an application fixed on one model, this requires timely maintenance, and the successor is in practice cheaper with superior quality.

What this means for design

The cost behaviour of an application over time depends much more on its architecture than on the market. Four principles are sufficient to place it on the price frontier, at a distance from a single provider's price list.
  1. We isolate the model call behind an abstraction layer, so that any model change becomes a configuration change.
  2. We maintain a business evaluation set, versioned, which qualifies a replacement model in a few days, to seize a price drop as well as to absorb depreciation.
  3. We measure the cost per task, at the granularity that distinguishes market deflation from the inflation of its own uses.
  4. We clarify the migration policy, because over the period studied, an equivalent quality model at least three times cheaper appeared every six to twelve months.
For an application designed according to these principles, with a constant scope, the inference cost decreases for approximately three years, then settles on a floor that the market holds at a few cents per million tokens, and no provider can decide otherwise, since both the open source floor and your ability to leave are beyond their control.

Sources: a16z, « Welcome to LLMflation » (2024) · Epoch AI, « LLM inference price trends » (2025) · Stanford HAI, AI Index 2025 · Gundlach et al., « The Price of Progress » (arXiv:2511.23455) · Xiao et al., « Densing Law of LLMs » (Nature Machine Intelligence) · SemiAnalysis (2026) · price lists and depreciation schedules from publishers. The "mixed" input and output prices are weighted three to one, according to Epoch AI's convention, and these are advertised prices, excluding discounts, therefore upper bounds. Points after February 2025 are reconstructed from public tariffs, as Epoch AI has not published beyond that date.
This study comes from Galadrim's Data & AI team. The agency's clients pay for their inference directly with the providers, without runtime intermediation, and the question comes up often enough to warrant a data-driven answer. The design principles above are those the team applies to its own projects.

Do you want support to launch your digital project?

Submit your project now