September 2026 · 6 min read
Uber’s 2026 AI budget
died in 4 months

Key Definitions
Consumption metering Pricing by tokens consumed rather than by fixed per-seat fee. Claude Code is not priced per seat; it meters tokens across model calls — the same engineer on the same workday can produce wildly different invoices depending on whether they run autocomplete or orchestrate parallel agents across a monorepo.
Per-seat pricing A flat monthly fee per user, e.g. Microsoft 365 Copilot Enterprise at $30/user/month. Finance can multiply headcount into a full-year line item; the vendor simply caps its upside on heavy users. Both models are defensible — treating them as interchangeable in a planning cycle is what produced Uber’s outcome.
Committed-spend agreement A procurement contract that locks a fixed amount and fixed rates in exchange for committed usage, trading volume for price predictability. After Uber, procurement teams that want predictability will negotiate these instead of riding consumption pricing.
When 84% of your engineers use the same coding agent, the budget question stops being “should we” and becomes “how do we account for it.” Uber exhausted its entire 2026 AI tools budget by April — the tool did not fail and engineers did not misuse it; they used it for exactly the workloads it was designed for: parallel agents, large-scale refactoring, automated test generation. Our judgment: the failure mode of coding agents is not rejection — it is ungoverned embrace, and consumption-priced tools need DevOps-grade cost control. Source: Forbes, company-reported.
The numbers behind the burn
Uber rolled out Claude Code in December 2025. Adoption climbed from 32% of engineers in February to 84% classified as agentic coding users by March; by spring, 95% of engineers used AI tools monthly, roughly 70% of committed code originated from those tools, and about 11% of live backend updates were written by agents with no human in the loop. CTO Praveen Neppalli Naga confirmed the overrun to The Information, saying the company was back to the drawing board on its assumptions. Uber’s R&D spend reached $3.4 billion in 2025, up 9% year over year — so this collapse is less about scale and more about a pricing model enterprise finance has not learned to manage.
The distribution of cost is the instructive part: monthly cost per engineer averaged $150–$250, with power users between $500 and $2,000; Naga himself reported spending $1,200 in a two-hour session. Engineers used the tool for exactly the workloads it was designed for — parallel agent execution, large-scale codebase refactoring, automated test generation, backend code production. From a productivity standpoint the rollout was a success; from a finance standpoint it was a runaway.
Why per-seat budgets cannot absorb token bills
Claude Code does not price per seat. It meters tokens consumed across model calls: an engineer running autocomplete consumes a fraction of what an engineer orchestrating parallel agents across a monorepo consumes. The same tool, the same engineer, the same workday can produce wildly different invoices depending on workflow choice. Annual budget cycles built around predictable per-license costs cannot absorb that variance.
Contrast Microsoft 365 Copilot Enterprise at $30 per user per month: a flat line item finance can multiply by headcount, with the vendor’s upside capped. Anthropic’s consumption model gives the vendor unlimited upside on heavy users and finance almost no forward visibility. Both models are defensible; treating them as interchangeable in a planning cycle is what produced Uber’s outcome. Uber compounded the dynamic by ranking engineers on internal leaderboards based on usage — turning token consumption into a competition. The teams driving adoption were not the teams managing spend, and that organizational gap was the load-bearing flaw.
The structural shift: subscription pricing cannot carry agentic workloads
Behind the single case is a pricing migration. On May 13, Anthropic announced that from June 15, agent tools and third-party harnesses on paid Claude subscriptions would be metered separately at full API rates; GitHub moved Copilot to a credit-based system on June 1. Analysts expect most vendors to introduce separate consumption pools for agents within 12–24 months — the vocabulary will vary (credits, requests, messages, compute units) but the direction is set: flat-rate inference for unbounded agentic workloads never survived the math, and vendors will pass the cost mechanics through to buyers.
Our judgment: the failure mode is ungoverned embrace
The industry’s standard response to consumption-cost stories is that AI pays for itself in productivity. Uber complicates that argument. The marginal productivity gain from a senior engineer running agentic workflows must clear a much higher token-cost hurdle than the gain from autocomplete; 5–20x increases in per-developer consumption are documented in agentic mode, and no public benchmark shows a matching multiplier on output value. Productivity savings also do not appear in the same line item as AI cost, so finance cannot net them out inside a quarterly review.
Our position: pilot economics do not predict scale economics. Pilots run on a few engineers using autocomplete; production runs on whole teams delegating multi-step workflows to agents — an order of magnitude more consumption. Industry data confirms the governance gap: only 43% of organizations have formal AI governance policies and only 21% have mature agentic governance. Most enterprises do not yet apply to AI tooling the spending controls DevOps routinely applies to cloud compute: per-engineer caps, real-time token monitoring, alerts before overrun rather than after. Uber deployed organization-wide without those controls, and the result was visible within a quarter.
Four controls you can copy
1. Treat tokens like cloud compute
Set per-engineer and per-team token budgets with real-time usage monitoring and overrun alerts — this is DevOps cost discipline transferred, not a new governance invention.
2. Remove or cap usage leaderboards
Any rollout that rewards usage without capping it is an unbounded liability by design. Incentivize quality metrics (defect rate, cycle time), not token volume.
3. Separate cost models by workflow class
Autocomplete and multi-agent orchestration have different cost curves. Model them separately so finance knows orchestration-type work costs 5–20x per engineer.
4. Negotiate committed-spend fixed rates
Procurement teams that want predictability negotiate committed-spend agreements; their leverage depends on whether engineering has usage caps in place at all. No caps, no leverage.
Uber is not slowing down: Naga plans to test OpenAI’s Codex alongside Claude Code, with a long-term vision where agent engineers handle coding, testing, and deployment while humans orchestrate. That direction is consistent across major engineering organizations. The open question for boards is not whether to deploy — it is whether finance has any visibility into what these tools will cost once engineers stop holding back.
Actions and the question for buyers
Actions: before any organization-wide coding-agent rollout, install the four controls (token budgets, no leaderboards, workflow-class cost models, committed rates); pilot cost is not scale cost — do not build a full-year budget from pilot numbers. Question for buyers: what is your heavy-user monthly cost per engineer, and can finance see it in real time? If you cannot answer, you are the next company that burns its budget by April.
OOMeta AI
OOMeta’s position: consumption-priced coding agents need cost control as a launch precondition, not a post-hoc fix — tokens are the new compute line item and FinOps discipline transfers directly. Uber’s lesson is organizational: rewarding usage without capping it converts a budget into an unbounded liability. We run our own agent pipelines the same way: per-task budgets, real-time metering, alert on overrun.
Book a diagnosticReferences: Forbes, Uber Burns Its 2026 AI Budget In Four Months On Claude Code (company-reported; Janakiram MSV, 2026-05-17) https://www.forbes.com/sites/janakirammsv/2026/05/17/uber-burns-its-2026-ai-budget-in-four-months-on-claude-code | The Information interview with CTO Praveen Neppalli Naga (cited by Forbes) https://www.theinformation.com/newsletters/applied-ai/uber-cto-shows-claude-code-can-blow-ai-budgets | Axios on Anthropic’s separate agent credit meter (2026-05-14) https://www.axios.com/2026/05/14/anthropic-claude-price-openai-tokens | SiliconANGLE on GitHub Copilot credits (2026-05-14) https://siliconangle.com/2026/05/14/anthropic-announces-programmatic-credit-pool-agentic-tool-use-rises/ | InfoWorld on industry consumption pooling https://www.infoworld.com/article/4171274/anthropic-puts-claude-agents-on-a-meter-across-its-subscriptions.html
FAQ
How much did Uber actually spend?+
Uber exhausted its full 2026 AI tools budget by April. Claude Code adoption climbed from 32% of engineers in February to 84% by March; roughly 70% of committed code originated from AI tools and about 11% of live backend updates were written by agents with no human in the loop. Average cost ran $150–$250 per engineer per month, power users $500–$2,000, and the CTO spent $1,200 in one two-hour demo. Source: Forbes, company-reported.
Why does per-seat budgeting fail with token pricing?+
Per-seat budgeting is headcount times a fixed price — predictable for a full year. Token pricing follows usage: the same tool, engineer, and workday produce wildly different invoices depending on workflow (autocomplete vs. parallel multi-agent refactoring). Annual budget cycles cannot absorb that variance; that is the structural reason Uber blew through its budget.
Why are internal usage leaderboards dangerous?+
Uber ranked engineers by Claude Code usage, turning token consumption into a competition. The teams driving adoption were not the teams managing spend — that organizational gap was the load-bearing flaw. Any rollout that rewards usage without capping it should be modeled as an unbounded liability until proven otherwise.
How is vendor pricing changing?+
On May 13, Anthropic announced that from June 15 paid Claude subscribers’ agent tools and third-party harnesses would be metered separately at full API rates; GitHub moved Copilot to a credit-based system on June 1. Analysts expect most vendors to introduce separate consumption pools for agents over the next 12–24 months — flat-rate inference for unbounded agentic workloads never survived the math.
What is wrong with the “productivity pays for itself” defense?+
Agentic mode multiplies per-developer consumption 5–20x, yet no public benchmark shows a matching multiplier on output value. Productivity gains also do not appear in the same line item as AI cost, so finance cannot net them out in a quarterly review. The two live on different ledgers and cannot offset each other.
Which controls should I copy?+
Four: ① per-engineer/per-team token budgets with real-time monitoring (treat tokens like cloud compute); ② remove or cap usage leaderboards; ③ model cost by workflow class (autocomplete vs. multi-agent orchestration); ④ negotiate committed-spend fixed-rate agreements — leverage depends on whether engineering has usage caps at all.
Related Articles
Where did Farmers’ 16.4M freed hours go?
Farmers cut routine servicing 35% across 8,000 agents, freeing 16.4M hours a year for sales. Our judgment: freed time is a liability until reallocated.
Inference outspends training. Buy switching power
Gartner: 2026 inference spend ($23.3B) beats training ($19B); agentic costs >5x by 2028. Cost = configuration, not model price. Buy switching, not tokens.
400 agent seats at $14: governed cost goes public
MRH Trowe: 400 agent seats at $14/seat/month, governance included. Our judgment: control-included cost is real; stack-free quotes underquote by design.
Adoption sprints, EBIT stalls: a scaling-ROI gap
McKinsey 2026: agent scaling jumped 27% to 40%; EBIT impact stuck at 37%. High performers rebuilt workflows (73% vs 25%) — conversion, not tech.