September 2026 · 7 min read
AI cost’s upstream is memory,
not the model price list

Key Definitions
Jevons paradox (AI cost context) When efficiency gains push unit prices down, total consumption can grow faster, so total spend rises anyway. A falling unit price is not a falling total cost; the two must be modelled separately.
HBM (high-bandwidth memory) High-bandwidth memory packaged with a GPU or accelerator. It sets how much model, context and concurrency a single machine can hold, and its supply and price feed directly into inference cost per unit of compute.
Feature spacing The spacing of critical storage-cell structures. Tighter spacing packs more cells per unit area and yields more dies per wafer, amortising the cost of each die. CXMT's fifth-generation platform is at 11.95nm.
Everyone watches token prices, but the price list that will set AI unit cost in 2027 is not held by model vendors. TrendForce estimates global DRAM industry revenue rose 59.5% quarter on quarter in 2Q26 to nearly $154.7 billion, with supply expansion still lagging demand growth. In the same window CXMT put its fifth-generation platform into mass production with a 11.95nm feature spacing. Our judgment: the downstream curve of AI cost is set by the supply cycle of upstream silicon — betting all cost governance on choosing a cheaper model treats one controllable variable as the whole set.
1. Demand fights a price war while supply issues price increases
TrendForce estimates global DRAM industry revenue rose 59.5% quarter on quarter in 2Q26, reaching nearly $154.73 billion. The three major suppliers posted: Samsung $60.98 billion (+63.4% QoQ, 39.4% share), SK hynix $38.59 billion (+37.9%, 24.9%) and Micron $36.0 billion (+65.5%, 23.3%). Second-tier suppliers rose faster still — Nanya +68.3%, Winbond +75.8%. Source: TrendForce press release.
TrendForce attributes the growth to AI server demand — expanding LLM training and AI inference lifting demand for high-density memory — alongside sharp rises in conventional DRAM contract prices. The phrase in the report’s own title matters: supply expansion continues to lag demand growth. This is not a cyclical glut; it is supply failing to keep up.
Our judgment: the two curves are separating. Token prices on the model side are falling; the silicon that supports the tokens is rising. Any AI cost forecast built only on token prices is missing half its inputs.
2. CXMT’s fifth-generation platform is a structural supply variable
In September 2026 CXMT announced at the World Manufacturing Convention in Hefei that its fifth-generation technology platform had entered mass production: critical storage-cell feature spacing down to 11.95nm using quadruple patterning, plus two 24Gb LPDDR5X products, each holding 50% more data than the previous generation, both already in mass production. Source: Reuters report.
CXMT says its process capability now matches the industry’s most advanced mass-production nodes and stresses that the new platform gives electronics manufacturers an additional source of supply. That sentence matters more to buyers than the parameters: global DRAM has long been a three-player market, and a fourth generational competitor means procurement structure finally has real substitutability.
Where restraint is required: yields, cost and ramp are all undisclosed, and matching the most advanced node is a vendor self-assessment. A generational claim is not available capacity; procurement should still be based on actual supply and validation data.
3. Why memory is the hard constraint on AI cost
Within inference cost, HBM and host DDR5 set how much model, how much context and how much concurrency a single machine can hold. When memory gets expensive and tight, three costs rise at once: cost per unit of compute (more machines for the same model), context cost (KV cache footprint), and time to add capacity (no supply, no capacity).
Put differently, a cheaper model lowers the price of a call but not the silicon that supports it. When demand grows faster than supply, the latter’s increase catches up with and can consume the former’s decline — which is how models getting cheaper and total AI spend getting larger can both be true.
Our judgment: the most commonly omitted first-order assumption in AI cost planning is the memory supply cycle. Most enterprise cost models contain only model price times usage, and omit supply constraint times capacity lead time.
4. What this means for enterprise procurement
First, extend the planning horizon to 12 to 18 months and model memory price curves and availability explicitly, instead of tracking only a quarterly-updated model price list. Second, treat memory supply the way you treat cloud vendors — multiple sources, committed volumes, and generational transitions written into architectural assumptions, because capacity and power differences within a generation change how much concurrency a single machine carries. Third, buyers in China gain a procurement option: a domestic source with generational competitiveness changes the bargaining structure and redundancy design.
Conversely, lower expectations for cost optimisation limited to model selection. The ceiling on that lever is set by the model price list, and the model price list is one line item of total cost.
5. Our judgment and a planning checklist
Our judgment: the first-order constraint on AI unit cost is the supply cycle of silicon. Betting all cost governance on the model layer means optimising a line item that is not your largest.
Put memory in the cost model
Alongside model price times usage, add a line for memory and VRAM supply and price assumptions, reviewed quarterly.
Assume a 12 to 18 month window
Capacity lead time is set by supply, not by budget; procurement cycles must lead the demand curve.
Multiple sources and committed volumes
Treat memory supply like cloud vendors: at least two substitutable sources, with committed volume on critical capacity.
Treat generational shifts as architecture variables
Capacity and power differences across generations change concurrency and cost per task; architectural assumptions must move with them.
Lower expectations on the model layer
Model selection remains necessary, but it is not the main battleground of cost governance.
Action: the smallest move this month — break open the model price times usage line in your AI cost model, add memory and supply assumptions, then answer one question: if memory prices rise another 20% next year, which line breaks first? The question to leave with: what share of your cost model is actually under your control?
OOMeta AI
AI cost governance has to start upstream in the supply chain — model tier, compute configuration, memory supply and capacity lead time are line items on the same sheet. We model cost per completed task rather than cost per token precisely so that upstream shifts become visible.
Book a diagnostic sessionReferences: TrendForce, “DRAM Industry Revenue Rises 59.5% QoQ in 2Q26 as Supply Expansion Continues to Lag Demand Growth”, 7 September 2026; revenue and share figures are TrendForce estimates. trendforce.com/presscenter/news/20260907-13219.html | Reuters, “China’s CXMT says new memory-chip platform enters mass production”, 19 September 2026; 11.95nm, quadruple patterning and the 50% per-die capacity gain are vendor claims. Reuters report (syndicated).
FAQ
Why is memory the upstream of AI cost?+
Inference cost depends on the model size, context length and concurrency a single machine can hold, and those are set by HBM and DDR5 capacity and availability. A cheaper model lowers the price per call; it does not lower the silicon that supports the call.
What happened in the DRAM market in 2Q26?+
TrendForce estimates industry revenue rose 59.5% QoQ to nearly $154.7 billion: Samsung $60.98B (+63.4%, 39.4% share), SK hynix $38.59B (+37.9%, 24.9%) and Micron $36.0B (+65.5%, 23.3%).
What does CXMT's fifth-generation platform change?+
Feature spacing down to 11.95nm with quadruple patterning, and two 24Gb LPDDR5X products each holding 50% more data than the previous generation, already in mass production. It adds a fourth supply source with generational competitiveness — though yields, costs and ramp are undisclosed.
Models keep getting cheaper, so why is total AI cost rising?+
Because demand grows faster than efficiency: unit prices fall while usage grows faster still, and the memory supporting that usage is tight and getting more expensive. The two curves compound into higher total spend.
How should enterprise cost planning change?+
Extend the horizon to 12 to 18 months, add a line for memory and VRAM supply and price assumptions alongside model price times usage, review it quarterly, and keep at least two substitutable sources.
Are these figures trustworthy?+
The revenue figures are TrendForce estimates. CXMT's 11.95nm and 50% capacity gain are vendor claims, not independently verified. A generational claim is not available capacity; procurement should still follow actual supply and validation data.
Related Articles
Model buying is two markets: volume vs money
Open-weight models run 56% of tokens, 14% of spend; 9/10 teams stepped Fable 5 down to Opus 5. Our judgment: buy by task tier, not average token price.
Routing savings aren't bought: ACL’26 benchmark
ACL’26 LLMRouterBench finds commercial routers fail to beat a simple baseline (arXiv 2601.07206). Our take: deterministic policy beats buying a router.
Inference outspends training. Buy switching power
Gartner: 2026 inference spend ($23.3B) beats training ($19B); agentic costs >5x by 2028. Cost = configuration, not model price. Buy switching, not tokens.
Agentic open weights: verify, don’t believe
Smaug's 15-20% gains and 10-100x savings are vendor claims. Open weights make them checkable: evaluate on your own traces, measure cost per completed task.