Bitcoin taught the world an uncomfortable lesson about energy: electricity is a local, perishable, politically captive commodity right up until someone finds a way to turn it into a portable digital claim. A megawatt-hour in Sichuan cannot be shipped to Frankfurt. A block reward mined with it can.
In 2026, a second industry has learned the same trick — at a far larger scale, and with the explicit backing of a state that has spent twenty years perfecting the art of commoditising other people’s high-margin businesses.
The product this time is the inference token: roughly three quarters of a word, the unit by which a language model’s output is metered, billed and sold. And the numbers on who supplies it have quietly inverted.
The crossover nobody announced
OpenRouter is the largest vendor-neutral model router in the world, moving north of 20 trillion tokens a week. It is one of the few places where you can watch what developers actually pick when no procurement committee is looking.
On OpenRouter data compiled by Bloomberg and Exponential View, the combined share of US frontier labs — Google, OpenAI, Anthropic — fell from roughly 70% of token volume in June 2025 to roughly 30% in June 2026. A joint OpenRouter–a16z study covering more than 100 trillion anonymised tokens tracks Chinese open-weight models going from 1.2% to nearly 30% inside a year. DeepSeek is now the single largest provider by volume at 16.3%, ahead of all three American labs individually.
The week of 9–15 February 2026 was the crossover: Chinese systems processed more tokens than American ones for the first time. The ratio now runs roughly three to one.
This is not a benchmark argument. Benchmarks are marketing. Routing data is revealed preference.
The price structure
| Model | Company | Input | Output |
|---|---|---|---|
| DeepSeek V4 Flash | DeepSeek | $0.14 | $0.28 |
| DeepSeek V4 Pro | DeepSeek | $0.435 | $0.87 |
| MiniMax M2.7 | MiniMax | $0.30 | $1.20 |
| Kimi K2.6 | Moonshot AI | $0.68 | $3.42 |
| GPT-5.6 Sol | OpenAI | $5 | $30 |
| Claude Opus | Anthropic | $5 | $25 |
Per million tokens, USD, July 2026.
DeepSeek V4 Pro runs roughly 34x cheaper than GPT-5.6 Sol on output. The budget tier undercuts its American equivalent by about 21x.
Level matters less than trajectory. Chinese labs cut API prices six times in the first half of 2026, three of those cuts declared permanent — structural repositioning, not launch promotions. And the line item almost nobody prices correctly: cache hits. Moonshot holds cached input at $0.07 per million, DeepSeek at $0.0036. If you are running an agent that resends the same system prompt thousands of times a day, that single number decides whether your product has a business model.
Why it is cheap
Two reasons, and only one of them is subsidy.
Architecture. The leading Chinese models are Mixture-of-Experts pushed to an extreme, activating a small fraction of total parameters per forward pass. MiniMax M2.5 carries 229 billion parameters and activates 10 billion. Independent testing cited by The China Academy put DeepSeek V3’s inference cost at roughly 36x below GPT-4o’s. That is unit-cost engineering, and it does not reverse when the subsidies stop.
Energy. Total electricity costs in China run around 40% below US levels. The cluster sits in China, draws on the Chinese grid, and delivers its output to a developer in Zürich or Singapore.
That second point is the one this audience should sit with. The electricity never left the grid. Its value did.
Where the mining parallel holds — and where it breaks
Anyone who watched hashrate migrate to stranded hydro in Sichuan, then to flared gas in Texas, already understands the trade: energy arbitrage laundered through a digital good that crosses borders at the speed of a packet. No customs declaration. No tariff line. No entry in the trade statistics.
Beijing has even coined the phrase — token 出海, “tokens going overseas.”
But the differences matter more than the symmetry.
Bitcoin’s hash is permissionless and perfectly fungible: a block is a block regardless of who found it, and no one can revoke it after the fact. An inference token is neither. It carries a vendor, a licence, a jurisdiction and a version number. The provider can deprecate the model, change the terms, or log the prompt. What looks like a commodity from the demand side is still a branded, revocable service from the supply side — unless you take the open weights and host them yourself, which is precisely the option Chinese labs offer and closed American labs do not.
That optionality, not price, is what is quietly winning European and Asian enterprise deals. Data stays inside the perimeter. Compliance teams can pin a version. No US-hosted endpoint sees the payload.
A regulator started speaking commodity
In March 2026 China’s National Data Administration formally standardised the Chinese word for token — 词元, cí yuán — with its director describing tokens as measurable, priceable and tradable. Domestic daily token calls, on the same agency’s figures, went from 100 billion in early 2024 to 140 trillion by March 2026.
When a state agency reaches for the vocabulary of commodities, the positioning has already been decided. Measurable, priceable, tradable is the precondition for spot pricing, then forward pricing, then everything crypto markets have been building infrastructure for since the first DePIN whitepaper. The decentralised compute thesis has spent three years short of demand-side proof. It just got a supply curve.
Volume moved. Revenue did not.
Here is where most takes stop too early. On OpenRouter, Anthropic holds around 12.3% of tokens processed but reportedly captures near 46% of platform revenue, because a premium token is worth many commodity tokens.
Two markets are forming in parallel, not one changing hands. The commodity lane — classification, extraction, chat backends, retrieval — is measured in billions of tokens, “good enough” wins, and cost per token is the only variable that moves. Chinese open-weight models are running away with it. The premium lane is long-horizon agents and complex engineering work where being wrong is expensive: a feature-sized coding task costs roughly $0.21 on DeepSeek V4 Pro against $3.80 on a frontier Western model, and there are plenty of contexts where that spread is a rounding error next to the cost of failure.
BYD’s unit volume tells you nothing about Porsche’s P&L. The open question is whether the premium lane stays wide enough — and the capability gap is closing faster than incumbents would like.
What to actually do with this
Stop treating model choice as a decision. It is a routing policy, defined by task class and revisited quarterly. Sending everything to the most expensive endpoint is the most common and most expensive mistake in production AI.
Price the cache, not the sticker. For agentic workloads, cached input pricing dominates total spend.
Treat below-cost pricing as a strategy, not an equilibrium. Solar and EVs show what follows: consolidation, exits, price stabilisation. Over 400 Chinese EV makers have shut down since 2018 while sector margins fell from 7.8% to 4.3%. Stress-test your unit economics at 3x today’s token price.
Read the licence like a risk officer. Distillation allegations against several Chinese labs are noted in USCC material. Merits aside, what a regulated business needs is licence terms, version pinning and traceability. Those are risk line items, not cost line items.
Run your own eval. A model tops a leaderboard because it is cheap, new, or used by one high-volume customer — not because it fits your workload.
For two years the framing was a capability gap, with export controls widening the lag. That was the wrong question. The gap that matters is who supplies the tokens the world consumes, and on the open layer that contest is already settled. This is what commoditisation looks like: the frontier becomes a feature while the good-enough tier eats the volume.
And the macro version, for anyone who has ever thought hard about where a kilowatt-hour goes: China has found a way to export deflation in a form that clears no customs, pays no tariff, and shows up in no trade statistic.
The full pricing tables, market-share methodology and source list are in the original long-form analysis: China Has Found Its Next Commodity to Export: The Token.

