The Token-Price Reckoning
GitHub / Microsoft / OpenAI / DeepSeek / Sequoia / FSB IndustryWhy the cheap-AI era is ending, why physics is the reason, and what to do in the window that's still open.
In a single week, three stories that look unrelated told the same story. GitHub moved 4.7 million paid Copilot subscribers off flat-rate plans and onto pay-per-token billing.[1] Reporting surfaced that Microsoft was canceling most of its internal Claude Code licenses — not because the tool failed, but because engineers adopted it so heavily that token bills outran the budget.[2] And reporting documented that a large share of the AI data-center capacity America promised for 2026–2027 hasn't broken ground, throttled by a power grid that takes years to build.[3] Read together, these aren't three news items. They are three readings off the same gauge — and the needle is moving in one direction.
The price you pay per token has been a marketing number, not an economic one. That gap is now closing.
What actually happened — the billing model flipped, industry-wide
On June 1, GitHub retired flat-rate Copilot subscriptions and switched its paid base to consumption pricing, where credits are drawn down against the input, output, and cached tokens of every interaction. Base seat prices didn't change — Pro at $10, Business at $19, Enterprise at $39 per user — but those numbers now describe an allowance, not a ceiling.[1,4] This wasn't a GitHub idiosyncrasy. OpenAI converted Codex from per-seat licensing to pay-as-you-go in April;[5] Anthropic has long metered Claude by the token. The pattern is the tell: when every major vendor abandons the same pricing model inside a single quarter, the model wasn't a choice. It was a subsidy they could no longer afford.
The Microsoft signal
The most instructive data point came from inside Microsoft. Its Experiences and Devices division — the org behind Windows, Office, Teams, and Surface — introduced Claude Code to engineers in December and began winding most licenses down by June 30.[2] Adoption had hit 84–95% monthly usage, with per-engineer API costs running $500 to $2,000 a month, and engineers reportedly preferred Anthropic's tool over their employer's own.[2] The decision was management's, and the reason was financial, not technical: under token billing, the more value the tool delivered, the faster the budget vanished. Uber reportedly burned its full-year 2026 AI budget in four months.[6] That is the structural trap of consumption pricing for genuinely useful agents — productivity and cost are the same curve.
The physical ceiling
Then the supply side. The four largest U.S. hyperscalers still intend to spend more than $650 billion on AI infrastructure in 2026 — and they physically cannot deploy it on schedule.[3,7] Of roughly 12 GW of U.S. data-center capacity announced for 2026, only about 5 GW is under construction; the rest is canceled or stalled.[3] The bottleneck has moved downstream of the chips entirely. Nvidia is shipping. What isn't shipping is the grid: high-voltage transformers that ran 24–30 month lead times before 2020 now stretch to five years.[8] Interconnection queues run three to five years in the densest markets.[9] Gartner projects power shortages will constrain 40% of AI data centers by 2027.[10] Electrical equipment is under 10% of a data center's cost and, right now, 100% of the bottleneck.[8]
The argument: demand compounds, supply is bolted to physics
Put the two halves together and the thesis writes itself. Demand for tokens is compounding — agentic workflows multiply the token count of every human interaction, and the tools that drive that consumption are the ones people actually want to use. Supply is gated by the slowest-moving inputs in the industrial economy: grid interconnection, transformer manufacturing, transmission permitting, water, and the rare-earth and copper supply chains underneath all of it. Compute scales on a monthly cadence. Substations scale on a decade. When demand is exponential and the binding constraint is a transformer with a five-year lead time, price is the only variable left to do the clearing.
The part most coverage misses: today's prices were never real
Here is the load-bearing fact. Independent analyses through early 2026 converge on the same uncomfortable conclusion — the frontier labs have been pricing inference below cost, deliberately, to capture share. One widely circulated audit estimated that current AI tools are subsidized by something like 90–98% of their true cost to serve.[11] OpenAI is reported to be losing well over a dollar for every dollar of inference revenue.[12] Sequoia's now-familiar framing said it plainly: the spend-to-revenue ratios don't close without dramatic price increases, dramatic cost reductions, or both.[13] The seat-based, all-you-can-eat plans of 2024–2025 were the free-sample phase. The move to metered billing is the moment the sample table gets cleared.
That analogy is the right mental model — with one critical asterisk. Ride-sharing's costs were labor and fuel, which don't hit a hard physical wall. AI's marginal cost is electricity and silicon flowing through infrastructure that does. So the normalization here isn't just "investors want margin now." It's that even if the labs wanted to keep prices low, the grid won't let them serve the demand at scale without bidding up the scarce input. Multiple forecasters now flag a 30–50% upward repricing of frontier API rates within 12–24 months as capital discipline returns.[14] The honest range is that wide because two forces are fighting.
The honest counter-current: deflation — and a grid that doesn't share our ceiling
A complete picture has to name the force pushing the other way, because it's genuinely powerful. Inference is getting radically more efficient — sparse-attention architectures cutting long-context compute 40–60%, new caching techniques, hardware pairings claiming 5x throughput over prior GPU baselines.[15] And the open-weight and low-cost providers are pricing aggressively: DeepSeek's V4 Flash sits around $0.14 / $0.28 per million tokens, an order of magnitude or two below the frontier models from the largest U.S. labs,[16] and it made a 75% price cut permanent — framed not as a promotion but as an efficiency gain passed through.[17] Pinterest reportedly hit frontier-class quality at a 90% cost reduction by post-training an open model on its own data.[18]
And there is a deeper, structural reason the cheapest tier sits where it does — one that doubles back on this entire thesis. The supply bottleneck I've described is a U.S. bottleneck. One provider pricing well below the largest U.S. labs, DeepSeek, operates on the other side of that asymmetry. China added 543 gigawatts of power capacity in 2024 alone — more than the United States has added in its entire history[21] — and in 2025 added over 430 gigawatts of wind and solar, more than half the world's new renewable capacity, much of it wired directly into its data-center strategy.[22] Where U.S. power demand was flat for two decades and left utilities flatfooted, China's grew roughly 8% a year on the back of continuous grid investment.[23] Goldman Sachs projects China will hold around 400 gigawatts of reserve power by 2030, more than triple the entire global data-center fleet's expected demand, while the U.S. reserve margin slips toward the threshold where grid reliability itself comes into question.[24] It is a point that U.S. AI executives themselves have publicly acknowledged.[22]
Be precise about what this does and doesn't mean. It is not a claim about which country is "winning" — the two systems have different strengths, and chip access, model quality, and data-center footprint cut the other way. It is a narrower, structural observation: on the one input this thesis identifies as the binding constraint — electricity and the grid to move it — the providers feeding the cheap tier are not bound by the same ceiling that's about to push frontier prices up in the U.S. That is why the low-cost floor is so durable. It also sets up the real decision for a buyer, which is not about geography at all but about where your data and your regulators sit — the subject of the third recommendation below.
So which wins — the physical ceiling pushing prices up, or efficiency and competition pushing them down? The resolution is the one worth internalizing: the floor for the cheap, commoditized tier keeps falling, while the ceiling for frontier, always-on, agentic capacity keeps rising and gets harder to guarantee. The market is bifurcating. Routine tokens get cheaper; the high-stakes, high-availability tokens you'll actually want for the workloads that matter get more expensive and more rationed. Planning for an average price misses the whole point — the variance is the strategy problem.
What this means for you
1. Move fast on experimentation — the window is closing, not opening. The cheapest, most permissive access to frontier capability you will ever get is probably the access you have right now. Subsidized pricing is a learning subsidy: someone else is paying for your organization to figure out where AI creates real value. Treat that as a clock, not a constant. The teams that run aggressive, structured experiments this year will have mapped their high-value use cases before repricing hits; the teams that wait for "mature" pricing will be learning the expensive way, on the expensive bill.
2. Build provider optionality before you need it — especially on the low-cost tier. A single-vendor frontier strategy is now a cost-concentration risk. The right architecture routes work by stakes: reserve premium frontier models for genuinely hard reasoning, and push classification, summarization, drafting, and routine generation to efficient low-cost providers — Writer and other enterprise-grade options among them. The eye-watering price gaps (frontier models can run 35–100x the cost of the cheapest serious alternatives at equivalent context[19]) only matter if you've built the plumbing to exploit them.
3. Match the provider to the data — cheap isn't cheap if it creates exposure you can't accept. Some of the lowest sticker prices in the market come from providers hosted outside your own regulatory jurisdiction, and the discount is real. The relevant cost, though, isn't just per-token — it includes where data is processed and retained, which laws govern it, what contractual and audit protections exist, whether the model can be self-hosted, and what uptime or latency commitments come with it.[16,20] This is a data-residency and concentration question, not a flag-waving one: a U.S. enterprise has to weigh routing sensitive data to a China-hosted API, a European one weighs U.S. providers against GDPR and data-sovereignty rules, and any firm should think twice before concentrating critical workloads on a single vendor in any jurisdiction. The disciplined move is to map each use case to a provider whose jurisdiction and terms fit the sensitivity of the data — favoring options that can be self-hosted or contractually fenced for anything regulated, and reserving the cheapest external APIs for non-sensitive, low-stakes work. (One caveat to note rather than weaponize: some low-cost models face unresolved, unproven allegations about how they were trained; treat that as a due-diligence item, not a verdict.)
And one habit to build into everything
The deeper shift is in how you evaluate any AI use case. For two years the question was "does it work?" The new question, attached to it permanently, is "what does it cost to run at scale — in tokens?" A workflow that's a clear win at subsidized rates can quietly become a budget hole once it's metered and once it scales, and agentic designs are the most exposed because they multiply token consumption per task. Context windows stuffed beyond need, premium models used where a cheap one would do, multi-step chains that re-process the same material — each compounds. Make downstream token economics a first-class evaluation criterion, not an afterthought. Measure outcomes, not tokens — but never stop counting the tokens.
The bottom line
| The force | What it does to your bill |
|---|---|
| Subsidy ending | Below-cost pricing normalizes upward; flat seats give way to metered tokens. Expect 30–50% frontier repricing over 12–24 months. |
| Physical bottleneck | Grid, transformers, and permitting cap how fast supply can meet demand. Five-year transformer lead times set the floor under premium capacity. |
| Efficiency + competition | Pushes the commodity tier cheaper. The market bifurcates: routine tokens fall, frontier/agentic tokens rise and get rationed. |
| Your move | Experiment now, build provider optionality, price every use case in tokens, and match each provider to the sensitivity of the data it handles. |
- [1] GitHub Blog, "GitHub Copilot is moving to usage-based billing," Apr. 27, 2026. github.blog
- [2] Dapta, "Microsoft Drops Claude Code Over Runaway AI Token Costs," May 2026. dapta.ai
- [3] TechSpot, "Nearly half of US data centers planned for 2026 are facing delays or cancellation," Apr. 4, 2026 (citing Sightline Climate). techspot.com
- [4] Notebookcheck, "GitHub Copilot drops flat-rate billing and developers are not pleased," Jun. 2, 2026. notebookcheck.net
- [5] Beri.net, "Flat-Fee AI Dies: Your $99/Month Just Became $900/Month," Apr. 2026. beri.net
- [6] Cybernews, "Microsoft drops Claude Code after budget burned," May 2026. cybernews.com
- [7] Tech-Insider, "U.S. AI Data Center Delays: 7 GW Capacity Crisis," May 2026. tech-insider.org
- [8] Tech Fund (techinvestments.io), "Power Bottlenecks & The AI Data Center," Apr. 2026. techinvestments.io
- [9] DataPowerSupply, "AI Data Center Power Bottleneck: 40% Face Shutdown by 2027," May 2026; PJM interconnection data via Data Center Knowledge, Jun. 2026. datacenterknowledge.com
- [10] Enki AI, "AI Data Center Grid Strain: Power Halts Growth in 2026" (citing Gartner), Apr. 8, 2026. enkiai.com
- [11] "State of AI 2026: The $600B inference subsidy, energy bottlenecks, and labor," discussion on Hacker News, Mar. 10, 2026. news.ycombinator.com
- [12] AI Insights, "OpenAI Is Losing $14 Billion in 2026 — And Your AI Bill May Be Next," Mar. 13, 2026. aiinsightsnews.net
- [13] MindStudio, "The Free Sample Phase: Why AI Tools Are Underpriced and What Comes Next" (citing Sequoia analysis), 2026. mindstudio.ai
- [14] Oplexa, "AI Inference Cost Crisis 2026: Why Your AI Bill Is Exploding," Mar. 31, 2026. oplexa.com
- [15] AI Automation Global, "AI Inference Cost Crisis 2026: Why OpenAI Loses $1.35 Per Dollar Earned," Mar. 16, 2026. aiautomationglobal.com
- [16] CloudZero, "DeepSeek pricing 2026: V4, R1, API costs, and how to optimize," Jun. 2026. cloudzero.com
- [17] The Next Web, "DeepSeek made its 75% discount permanent. The AI price war just escalated," May 2026. thenextweb.com
- [18] VentureBeat, "How DeepSeek's radical architecture is shattering Silicon Valley's token moat," May 2026. venturebeat.com
- [19] TechJack Solutions, "DeepSeek Pricing 2026: The Cheapest AI API Costs Revealed," May 2026. techjacksolutions.com
- [20] InfoWorld / Computerworld, "DeepSeek's steep V4-Pro price cut escalates AI pricing war," May 2026. infoworld.com
- [21] IEEE ComSoc Technology Blog, "China vs. U.S.: Race to Generate Power for AI Data Centers as Electricity Demand Soars," Feb. 16, 2026. techblog.comsoc.org
- [22] Al Jazeera, "China's secret weapon in AI race with US? Lots of cheap energy," May 28, 2026 (citing IEA and Data Center Watch). aljazeera.com
- [23] Wood Mackenzie, "Powering China's data centres," Jul. 2025. woodmac.com
- [24] Goldman Sachs research summary via Tiger Brokers, "Power Supply Emerges as Key Variable in AI Infrastructure; China's Abundant Electricity Gives Competitive Edge Over US," 2026. itiger.com