Kevin Blackman. Published 2026-08-23.

Demand for Intelligence Is Insatiable - Part Two

Demand for Which Tier? Cheap Tokens and Frontier Tokens Just Moved Opposite Ways.


Two doors in one building - a crowd at the cheap entrance, a short unqueued line walking into the frontier one
Two doors in one building - a crowd at the cheap entrance, a short unqueued line walking into the frontier one

The short version

Everyone in this industry will tell you demand for AI is effectively unlimited, and everyone who says it is long the answer. Part One went looking for measurements instead of opinions and found quantities: nine hundred million people a week on the assistant, and plausibly at least tens of millions -- on the agent, overwhelmingly programmers, among whom adoption is already finished. No vendor publishes a category total for agents, and programmers are about 8% of knowledge work. A big runway, but a runway measured in users.

This article measures the other half, which is a good deal harder to spin. A user count tells you how many people showed up, while a price tells you what they would pay to stay.

The prices say that the cost of standing at the frontier has doubled in a year, and in the week to the 21st of August 2026 the second most expensive model in the market cut its price by nearly thirty percent. Meanwhile the cheapest way to buy the cheapest model has not risen at all -- it is about four cents a million tokens and still falling.

That is not what insatiable demand looks like, and it is not what collapsing demand looks like either. What it does look like is two goods with different supply and demand conditions that happen to share a name.


The month buyers started trading down

Four things happened within about a month in mid-2026, three of them decisions that somebody took with real money at stake and one a market repricing itself in public.

Four moves in one month
Four moves in one month

None of this shows that demand is weak, but it does show demand behaving differently at the two ends of the market, because at the frontier the price went up and buyers moved work down to cheaper models, while at the cheap end the price kept falling, since anyone with a data centre can serve an open-weight model.

The industry talks constantly about being sold out, which is a flattering problem to have, and says far less about the frontier customers who looked at the price and quietly moved down a tier. To call demand insatiable is to say that price does not change what buyers do, and at the frontier it plainly does.


The cost of the frontier has doubled in a year

The Token Price Index (TPI) tracks what it costs to use the best available AI, by averaging the prices of 21 leading models from 10 providers. We read it on the 21st of August and rebuilt its arithmetic from its own published inputs, matching its published figure exactly.

The blended price of frontier tokens came out at approximately $2.35 per million tokens, up 99.6% year on year. That is to say, the cost of standing at the frontier has almost exactly doubled, over precisely the period in which everyone was being told inference costs fall tenfold annually.

The cost of standing at the frontier - the year, and the six published weeks
The cost of standing at the frontier - the year, and the six published weeks

Both things are true at once, and they are not in conflict. Any given level of ability gets cheaper every year, which is the most impressive thing about this industry, while standing at the front of the queue gets dearer, because the front keeps moving and the newest models cost more than the ones they replace. Hence the rule we now apply to every claim about falling costs:

Never write "inference costs" without saying cost of what.

It also means that asking whether demand for intelligence is insatiable is the wrong question, because cheap tokens and frontier tokens are two different products whose prices have been moving in opposite directions. The question worth asking is "demand for which tier?", in the same way that asking whether demand for transport is insatiable means little until you say whether you are taking the bus or an Uber. The bus is cheap and crowded, runs to somebody else's timetable, and is the one that fills up and runs out. The Uber costs more, comes to you, and is nearly always there.


The index floor is not the market floor

Neo-cloud providers offer open-weight models at widely varying prices. TPI primarily uses the list price that the open-weight model's original author quotes at. This is nowhere near the real market price that this model is offered at when its multiple providers are considered. We therefore survey the real market prices across as many providers as we can, using TPI's blend formula, and found the following:

Where you buy DeepSeek V4 Flash Blended $/M
StreamLake, Baidu Qianfan $0.078
DigitalOcean $0.098
GMICloud $0.109
DeepInfra, Sail Research $0.117
CoreWeave, Novita, Parasail, AtlasCloud $0.182
Azure (US) $0.315
DeepSeek's own API, off-peak $0.352
DeepSeek's own API, peak $0.704
Eighteen ways to buy the same model
Eighteen ways to buy the same model

TPI books the bold row, which is the price DeepSeek quotes rather than the price the model sells for. So when DeepSeek raised its own prices in August, the cheapest model in the index nearly doubled from $0.182 to $0.352, and the index recorded that frontier tokens had become more expensive when nothing of the sort had happened to the market. A single seller had moved its list price, and an index built on list prices duly wrote the move down, while the same model remained available elsewhere at $0.078 and continued to fall.

A separate price history of eighteen models, each sold by several providers, matches our survey to within one percentage point and shows that until late June the authors were still the cheapest place to buy their own models. What neither can record is how a provider actually runs the model, so a cheap listing may be a compressed and slightly worse version sold under the same name, and that is a possibility we cannot yet rule out.

Why would you cut the price of your best product?

There are two plausible explanations for OpenAI cutting the price of GPT-5.6 Sol, its second most expensive model, and they could both be true at once. The duller one is that the capacity is already bought, since Microsoft, OpenAI's principal compute supplier, carries $443.5bn of operating and finance leases in its own filings and booked $24.1bn of revenue from OpenAI in FY2026 as a related party, and capacity on that scale is paid for whether or not anybody uses it. Once the money is spent and the lease signed, filling the machine at a lower price beats not filling it.

The other is that Anthropic has filed to go public, with a confidential draft Form S-1 on the 1st of June and Bloomberg reporting on the 15th of July that its banks were scheduling investor meetings for a Nasdaq listing as soon as October. An IPO is priced on growth and market share rather than gross margin, so weakening a competitor's growth figures before the roadshow can be worth more than the revenue you give up doing it. Against the first explanation sits one sentence at the bottom of OpenAI's own pricing page:

"GPT-5.6 Sol's promotional pricing is available at least through November 21, 2026."

Money already spent argues for cutting the price and leaving it cut, since a data centre lease does not expire on the 21st of November, whereas a price with an end date on it is the shape of a campaign. Nor is Sol the only one, because Gemini 3.7 Flash is on an introductory rate that doubles on the 1st of January, so if both revert as published then the cost of frontier access rises without a single provider deciding anything. We do not claim to know why OpenAI cut the price, and we would not believe anyone who said they did, but the end dates are facts rather than interpretations and they sit more comfortably with one reading than the other.

What we cannot see, and the real bear case

Four things our instruments would miss, the first of which is the strongest argument against everything above.

Open-weight models running on your own hardware. Asked for the case against all of this, the investor Gavin Baker did not reach for financing or valuation but for demand destruction, calling edge AI "by far the most plausible and scariest bear case" and putting it three years out, when a phone with enough memory will run a pruned-down Gemini or ChatGPT at thirty to sixty tokens a second, for free. It is already here on the desktop, because Qwen3.8-27B is free, runs in 17GB, and scores level with a hosted OpenAI model on a widely used capability ranking, so the same standard of work now costs either about fifty cents a million tokens or nothing at all. Work that moves onto your own machine has not disappeared, it has left the market we measure, and on the way out it looks exactly like a price decline. Demand for AI tokens will not necessarily fall. What is impacted is delivery from the cloud, so if that continues then token use keeps rising while the revenue from serving them falls, and it would hit the tier we currently read as the healthy one.

Rationing by degradation. Tight capacity does not always arrive as a price rise, since you can instead be throttled or quietly served a smaller version of the model, and a shortage that shows up as worse quality never reaches a price series at all.

Work with a right answer. Everything here is priced against office work, and we have no series for mathematics, chemistry, biology, materials or robotics, which are the fields where an answer can be checked and where being right is worth far more than the tokens spent getting there.

What the output is worth. We measure what intelligence costs to produce, while the industry's claim rests on the gap between that cost and its value, and the value side is barely measured at all. The most serious attempt we know of is SemiAnalysis's model of revenue per megawatt of datacentre capacity, which is solid work and sits behind a paywall we have not paid, so we can read its headline figures and not its method. A token is also worth more once the model behind it has been tuned to a particular business and pointed at that business's own tasks, and enterprises have barely started doing this, mostly on open-weight models because they can be modified, and they are cheap.


A tenfold price cut bought 13x usage, then stopped

Part One gave the bull case its due and it survived most of the way, and what is left pulls in two directions, with the part we cannot answer coming first.

Etched put it this way: tokens today are "handcrafted, like they made screws back in the Renaissance." If that is right then everything above is a pause in prices rather than a ceiling on demand, and we are nowhere near the economies of scale that phones and cars eventually found. There is a real measurement behind the idea, too, because OpenRouter cut the delivered price of one model tenfold, with OpenAI taking 5x off the list price and OpenRouter a further 2x off its own margin, and usage then rose 13x, which is more than the price cut alone would explain.

But read the next sentence of that account, because it is the one nobody quotes. Usage, in OpenRouter's own words, "grew and flattened out at 13x", and then resumed growing at roughly its old rate. It settled at a new level and stopped there, which is not insatiable demand but demand that responds to price and then runs out, so that at a given level of ability a tenfold price cut buys somewhat more than tenfold the usage and nothing further until the ability itself improves.

Their CEO is careful about it, and we will be too: "no one has done a good job modelling it, but there are spot stories that confirm it." One product, one price move, one platform.

What that measured was not whether demand for intelligence is insatiable, which nothing can measure, but something narrower and a good deal more useful, namely how much extra usage a given price cut buys while the ability of the model stays the same, and the answer turned out to be large, real and finite.

We cannot rule out another order of magnitude of cost decline, but we can say what happened the last time one arrived, and it was not insatiable.

How many more cost declines are left

There is no result that would settle the question of whether demand for intelligence is insatiable, because any evidence against it can be answered by pointing at some use that has not arrived yet.

However there is a version of this question that we can attempt to answer:

How many more order-of-magnitude cost declines are left, and how fast do they arrive?

That is a physics and manufacturing question with observable inputs:

On the last of those: if doubling the money spent produces less than double the output, the cost per token stops falling no matter how much is invested. If it is linear or better, costs keep dropping as the buildout scales.

Each of those has an answer that can be measured, and the answers would change what this sector is plausibly worth, though honest figures across the supplier base remain very hard to come by.

Thus what we are really watching:

The third tripwire carries a limitation we should name rather than let a reader find. Evan Conrad, who runs the GPU market at SF Compute, says that on the current generation of chips the large clusters are no longer sold on demand at all: the customer signs a long-term contract before the cluster is built, because the last cycle wiped out everyone who built on spec. What is left on the spot and on-demand market is the smaller deployments. So this instrument still measures a real market, but a shrinking corner of one, and it is watching the tier furthest from where frontier capacity is actually procured. If it stays quiet, that is weaker evidence than it looks.

Three tripwires, whose data are all fairly public, and are thus checkable. We have told you where ours are; ask the people telling you demand is insatiable where theirs are.


Every figure traces to a primary source: rate cards and labour statistics we read ourselves, index arithmetic we reproduced ourselves, and named practitioners on the record. Where a source is positioned in the sector, we have said so. Where we have guessed at motive, we have labelled it a guess. Corrections will be logged below this line.

Post-publication log

2026-08-23 -- limitation added to the third tripwire. The spot-versus-on-demand GPU tripwire was published without noting that on current-generation chips the large clusters are no longer sold on demand at all; buyers contract before the cluster is built. The instrument and the figure are unchanged. What changed is our claim about what a quiet reading proves: it is weaker evidence than we implied, because it watches a shrinking corner of the market. Source is Evan Conrad of SF Compute, who is an interested party and is cited as such.