September 2026 · 5 min read
Model buying is two markets:
volume vs money

Key Definitions
Model tier step-down Moving a production workload from the most capable (and most expensive) frontier model to the tier below when the capability delta does not justify the price delta. Vercel measured it: Fable 5’s gateway spend share fell from 13.2% to 4.9% while Opus 5 rose to 22.5% — nine in ten teams that ran Fable cut their usage.
Volume–spend split The structural gap between which models carry the most tokens and which carry the most cost. In August 2026 open-weight models ran 56% of AI Gateway tokens but took only ~14% of spend; spend stayed concentrated on frontier models.
Average token price Estimated spend divided by token volume, computed from labs’ published list prices on gateway telemetry. It fell 23.2% in August, the third straight monthly drop, and has halved in five months — a headline that hides the volume/spend split underneath.
The cheapest model is not the winner — the model that fits the task is. Vercel’s September AI Gateway Production Index is the first neutral, public time-series of what production AI actually costs, and it shows a structural split: open-weight models ran 56% of gateway tokens in August but took only ~14% of spend; nine in ten teams that used Anthropic’s top model stepped down to the tier below; and the average token price halved in five months. Our judgment: “average token price” is the wrong unit of purchase. The decision unit is model tier per task — and the numbers to trust are the ones you measure on your own workloads.
1. Volume and money are splitting into two markets
In August 2026, open-weight models ran a majority of tokens on Vercel AI Gateway for the first time: 56%, up from 7% in December 2025 — eight months from under one in ten to more than all closed-weight models combined. But volume is not money. The frontier kept the majority of spend, and open-weight dollar share, while accelerating, is still a minority. Anthropic alone has taken at least 61 cents of every dollar spent through the gateway every month since December, and 64 cents in August. Source: Vercel’s report.
What this shows is not “open source is cheap” — it is that two markets are forming: a commodity volume layer (open weights) and a quality-premium money layer (frontier). For C-suite the implication is direct: the “we went open weight so we saved money” narrative is wrong — the money is still in the frontier. The real question is whether each task tier’s spend is justified by its output.
2. Nine in ten teams stepped Fable 5 down to Opus 5
The cleanest proof comes from inside one lab. When the US export control on Fable 5 was lifted on July 1 and access restored, its gateway spend share surged to 13.2%. At the end of that same month Opus 5 came online. In August, Fable 5’s spend share fell to 4.9% while Opus 5’s rose to 22.5% — nine in ten teams that ran Fable cut their usage, and more moved to Opus 5 than to any other model. Vercel’s conclusion in one line: Fable’s extra capability wasn’t worth double the price.
The detail worth noting: the workload stepped down, but the money did not leave Anthropic. Teams went to Opus 5, not to another lab. That is the second pattern — customer loyalty follows the model profile, not the brand: a new model that preserves what users valued in its predecessor keeps them; one that doesn’t loses them to other providers.
3. Average token price: read the trend, not the number
The average token price fell 23.2% month over month in August — the third straight drop and the steepest since April — and costs less than half what it did five months ago. Among teams running more than ten million tokens in both months, the median team paid 7.6% less. For budgets, this is the good trend: the same budget buys more inference, and frontier models can be reserved for the tiers that justify the premium.
But the single number misleads. First, the methodology: spending is estimated at labs’ published list prices, actual bills may differ, and the report notes methodology is revised. Second, the average hides the structure: Google’s token share fell from 30% to 5%, with 22 of the 25 percentage points coming from Gemini 3 Flash alone — half its volume went to cheaper models like GPT-5.6 Luna, most of the rest to Opus 5 and Sonnet 5 at roughly nine and three times the price. The average barely moved; the decisions changed completely.
4. Our judgment: the purchasing unit is model tier per task
“Average token price” invites the wrong decision: pick the cheapest. But the 9-in-10 Fable-to-Opus step-down shows real buyers already optimize per task — they want “good enough”, not “most expensive” and not “cheapest”. Cost control was never about choosing the cheapest model; it is about choosing the right tier — reserve the frontier budget for loads that can prove the premium, put commodity volume on open weights, and back it with a quality gate.
Our position: delete “average token price” from your budget file and replace it with “cost per completed task by tier”. This is not a wording dispute — the 56%/14% split means “we went open weight so we saved” and “we picked the right tier for each task” are different claims, and only the second is real cost control. It is also how we run our own stack: route by task tier, measure by completed task, not by token.
5. The measurement caveat is the point
Vercel is the closest thing to a neutral observation point in public data, but it is still a gateway operator, and the report itself states: spending is a list-price estimate, actual bills differ, methodology is revised, and the comparisons track relative movement among the same teams rather than individual tokens. These caveats are not flaws — they are the argument: vendor-adjacent telemetry is not an audit.
For AI app leads: instrument per-task cost and quality in every tier; that is the routing signal. For CFOs: read “average token price” as a trend indicator, not a budget line. The bill-level numbers you can trust are the ones you measure on your own traffic.
6. Buyer checklist and the question left for you
1. List your task tiers
Chat, RAG, agent loop, batch — and a quality gate per tier with an explicit “good enough” bar.
2. Measure cost per completed task per tier
On real traffic, not list-price averages; this is the shared input for both routing and budget.
3. Re-check tier assignment quarterly
The Fable-to-Opus step-down happened in one month; the model landscape moves faster than budgets, and your tier map must move with it.
Action: reserve frontier spend for loads that can prove the premium; put commodity volume on open weights behind a quality gate; and start building task-level cost telemetry in your stack today. The question left for buyers: if you replaced every “average token price” in your budget deck with “cost per completed task by tier”, would any decision change? If not, you are not buying by fit yet.
OOMeta AI
OOMeta’s position and practice: route by task tier and measure cost per completed task, not per token — tier-level cost telemetry is the shared foundation for routing and budget. Vercel’s index is the closest public neutral series, and its list-price caveat is exactly why real cost audit has to come from your own workloads.
Book a diagnostic callReferences: Vercel, AI Gateway Production Index — September 2026 (spending estimated at list prices; methodology caveat stated) https://vercel.com/blog/ai-gateway-production-index-september-2026 | The New Stack, open-weight volume vs Anthropic spend https://thenewstack.io/open-weight-anthropic-spend/
FAQ
Where does the 56%/14% split come from?+
Vercel AI Gateway production telemetry through August 2026: open-weight models ran 56% of all gateway tokens (up from 7% in December 2025), but open-weight dollar share is still a minority — Anthropic alone kept 64% of spend. Volume and money are splitting into two markets.
Why did nine in ten teams leave Fable 5?+
When Opus 5 launched at roughly half the price, nine in ten teams that ran Fable cut usage and moved workloads to Opus 5 more than to any other model. Vercel’s conclusion: Fable’s extra capability wasn’t worth double the price. The spend stayed with Anthropic.
Can average token price be used as a benchmark?+
Only directionally. Spending is estimated from published list prices, actual bills differ, and methodology is revised — Vercel states this in the report. Use it to read trends (prices halved in five months), not as an audit of your own bill.
What happened to Gemini 3 Flash?+
It lost 95% of its gateway token share since May; more than three-quarters of that volume moved to models from other labs — about half to cheaper models like GPT-5.6 Luna, most of the rest to higher-priced Opus 5 and Sonnet 5.
What is the September special report?+
OpenAI’s GPT-6 Astra took a third of OpenAI’s gateway spend within two days of launch and 7.7% of all gateway spend in its first twelve days — twice Fable 5.1’s 3.7% at the same price.
What should buyers take away?+
Stop comparing vendors by average token price. Define the task tiers in your stack, measure cost per completed task per tier, and route to the cheapest tier that passes your quality gate.
Related Articles
Routing savings aren’t bought: ACL’26 benchmark
ACL’26 LLMRouterBench finds commercial routers fail to beat a simple baseline (arXiv 2601.07206). Our take: deterministic policy beats buying a router.
Agentic open weights: verify, don’t believe
Smaug’s 15-20% gains and 10-100x savings are vendor claims. Open weights make them checkable: evaluate on your own traces, measure cost per completed task.
LLM routing: conditions behind 40-80% savings
Route each request to the cheapest capable model. Behind 40-80% savings claims: cheap share past 50%, an eval gate, closed-loop distillation.
Inference outspends training. Buy switching power
Gartner: 2026 inference spend ($23.3B) beats training ($19B); agentic costs >5x by 2028. Cost = configuration, not model price. Buy switching, not tokens.