Frontier Capability Cost
14 note(s)
Note
The buildout is justified by the premium that the most capable models command. So it is worth
asking what that premium actually is. Each point below is the cheapest model available at
or above its capability level — the efficient frontier of the market as it is priced today.
Cost rises with capability throughout, and then accelerates sharply near the top: the
steepest single step on the curve buys the last stretch of measured capability.
Note
The capability/cost frontier, September 2026. Capability is the Epoch Capabilities
Index; cost is list price per million tokens, blended 3:1 input:output, on a logarithmic axis.
Open-weight models are marked separately.
R76
Note
Read the right-hand column. Three quarters of the entire capability range costs under
ten cents per million tokens — 1.9x, 1.9x, 1.7x across the whole of it. Then the
curve does not bend, it breaks: 106x across the next fifteen percentiles. In
absolute terms, index 154.5 costs $0.094 and index 162.6 costs $10.00: a 5.2% gain in
measured capability costs a hundred and six times as much.
Revised 9 September 2026. Two things moved at once and they should not be
confused. Epoch refits its index as benchmarks are added and retired, so every score on
this panel changed together — a scale change, not a correction. Separately, the
price basis changed: the panel now uses the cheapest place a model can actually
be bought, and a vendor's own list price only where that vendor is the only source. The earlier
version of this note read “1.9x, then 3.5x, then 4.5x… the 75th-to-90th band costs
10x.” The break is in the same place and is far sharper than published.
Note
The composition of the frontier splits just as sharply, and along the same line.
All five frontier models below the break at index 156 are open-weight or Chinese. All
five above it are closed and American. Not a tendency — every single one.
The cheap frontier is largely open; the expensive frontier is entirely closed and
American. That is consistent with the separate finding that seven of the ten most-used models on
OpenRouter are open-weight
R56.
Note
One of these prices has an expiry date printed on it. Gemini 3.7 Flash sits on the
frontier at $1.50 blended, but Google's own pricing page states that rate holds only through
31 December 2026, rising to double on 1 January 2027
R77.
At the new rate it leaves the frontier entirely.
A frontier partly composed of introductory
pricing is not a stable frontier — and an aggregated price feed reports the current
number with no expiry attached. This one was visible only by reading the vendor's page.
Note
Capability: Epoch AI, "AI Benchmarking Hub", published online at epoch.ai, licensed CC-BY.
Cost: each vendor's own published pricing page, read 20 August 2026, with an MIT-licensed
aggregated price map used as a cross-check. Where the two disagreed by more than 5% the vendor
page was used — this happened once, on GPT-5 nano, where the aggregator was 9.1% high.
Note
The cost axis is price per token, not price per task. A verbose reasoning
model emits more tokens to answer the same question, so per-token price understates its true
cost per task — and understates it most at the top of the range, which is exactly where
this chart is most interesting. The direction of that bias is known; its size is not.
Note
Capability indices are constructions, not measurements. The index used here
is a composite over one organisation's benchmark suite. A different suite would move the
points; the question is whether it would move the shape.
Note
200 of 283 indexed models carry a published price. Every current-generation
unpriced model in the band where it could displace a frontier point was priced by hand
against its vendor's page, because a missing price silently drops a model out of the
frontier — and dropping a cheap capable model would exaggerate the premium in the
direction this page already argues. The remainder are superseded previews and duplicate
host listings.
Note
Two different kinds of price sit on this chart. Where a model's creator
sells it, the creator's published list price is used. Where the creator only released
weights and runs no API of its own — gpt-oss, Gemma, Mistral NeMo, Qwen2.5-Coder
— there is no such price, so the cheapest third-party hosted rate is used instead.
Hosted rates for a single model span up to 19x across providers; the minimum is taken
deliberately, because the frontier asks what is cheapest available.
Note
The capability index does not cover every lab. Tencent, among others, has
no entry in it at all, so its models cannot appear here however cheap or capable they are.
A benchmark suite that omits a vendor removes that vendor's points from the cheap end of
the frontier, which flatters the premium rather than understating it.
Note
This prices cloud inference, and only cloud inference. Every figure here
is what somebody charges to run a model on their machines. Work that migrates onto the
buyer’s own hardware — an open-weight model on a workstation, or in time a phone
— leaves this series entirely, and its departure looks identical on the chart to a
fall in the cost of production. Gavin Baker, asked for the case against the buildout,
volunteered exactly this: edge AI is “by far the most plausible and scariest bear
case”. We would measure a deflation and report a cost curve, and nothing in this
panel would distinguish the two.
Note
List prices only. No batch discount, no cache discount, no negotiated
enterprise rate. Large buyers do not pay these numbers.
Note
- The cost axis is price per token, not price per task. A verbose reasoning
model emits more tokens to answer the same question, so per-token price understates its true
cost per task — and understates it most at the top of the range, which is exactly where
this chart is most interesting. more →
- List prices only. No batch discount, no cache discount, no negotiated
enterprise rate. more →
- Capability indices are constructions, not measurements. The index used here
is a composite over one organisation's benchmark suite. more →
- 200 of 283 indexed models carry a published price. Every current-generation
unpriced model in the band where it could displace a frontier point was priced by hand
against its vendor's page, because a missing price silently drops a model out of the
frontier — and dropping a cheap capable model would exaggerate the premium in the
direction this page already argues. more →
- Two different kinds of price sit on this chart. Where a model's creator
sells it, the creator's published list price is used. more →
- The capability index does not cover every lab. Tencent, among others, has
no entry in it at all, so its models cannot appear here however cheap or capable they are. more →
- This prices cloud inference, and only cloud inference. Every figure here
is what somebody charges to run a model on their machines. more →