In August 2026, OpenRouter routed 379 trillion tokens, 17x what it routed in October 2025. DeepSeek carried 22.4% of them, OpenAI 10.5%, Tencent 10.3%, Xiaomi 8.5% and Google 7.4%. On our method the month came to about 84,000 tonnes of CO2e, give or take 28%. This post has the curve, the shares month by month, the fifteen most used models, and a calculator for cutting the number.
Market share for language models is usually reported from surveys, which ask companies which vendor they use. That tells you who bought what and very little about how much work gets done, because a company running a billion tokens a day and one running a thousand count the same. Tokens are the unit the models are actually paid in, and OpenRouter publishes them, per model, every day.
Why tokens, and why OpenRouter
OpenRouter is a gateway. A developer sends a request to one endpoint and it goes to whichever provider they chose, so OpenRouter can see which models get used, and it publishes the counts on its rankings page and through a dataset behind it. We sum that dataset by month and by provider, and we run every month through the same method the product uses to turn tokens into CO2e. The numbers below are the sums as of September 8, 2026.
It is a sample rather than the whole market. Traffic that goes straight to OpenAI's or Anthropic's own API never passes through OpenRouter, and neither does anything inside ChatGPT, Gemini or Copilot. The developers who route through a gateway are more price-sensitive than average, and some of the largest applications on it are chat and role-play products. That skews the shares towards cheap and free models. Even so, it is the largest public dataset that counts tokens per model, and its direction of travel is hard to argue with.
The year in one curve
Monthly tokens on OpenRouter, from the product's public endpoint. Fetched 13 Sep 2026, 17:19 UTC.
Monthly tokens went from 22.3 trillion in October 2025 to 379.4 trillion in August 2026. That is 17x in ten steps, a doubling roughly every 2.5 months, or about 33% growth a month. The 11 complete months shown add up to 1.28 quadrillion tokens. On our factors that is 282,000 tonnes of CO2e, which is what about 61,000 cars emit in a year, from one gateway.
August 2026 alone was 84,000 tonnes, with a band of 60,000 to 107,000, and about 174 gigawatt-hours of electricity at the data centre meter once host power, cluster utilisation and cooling are counted.
Who gained share, and who lost it
- DeepSeek22.4%
- OpenAI10.5%
- Tencent10.3%
- Xiaomi8.5%
- Google7.4%
- Z-AI6.7%
- Other34.3%
Google's share shrank. It was the largest named provider through the winter, peaking at 23.7% in January 2026 on Gemini 2.5 Flash, and it ended August 2026 at 7.4%. Its volume still grew, because everything grew, but its slice fell to a third of what it was.
DeepSeek's growth came in one step. Its share sat between 5% and 8% for 6 months, then V4 Flash arrived in April 2026 and the share went to 15% in May, 17% in June and 22.4% in August. Two builds of that one model now carry 18.6% of all traffic on the gateway.
Two providers that were not on the board in March 2026 are on it now. Tencent's HY3 went from nothing in April 2026 to 11% in May and has held around 10% since. Xiaomi's MiMo V2.5 climbed to 14.8% in July 2026 before settling at 8.5%. Together with Z-AI's GLM models, the Chinese labs held 48% of August 2026's tokens.
OpenAI held close to 10% all year, which means its volume rose 17x along with the market. GPT 5.6 Luna is its most used model at 6.6%. Anthropic is the counter-example: Claude 4.5 Sonnet was the second most used model on the gateway in October 2025 at 10%, and Claude Opus 5 sits at 1.9% now. That is a statement about who routes Claude through OpenRouter, not about Claude's use in general.
The other change is concentration. The six largest providers carried 38% of tokens in October 2025 and 66% in August 2026. The "other" bucket, which was 62% of the gateway a year ago, is now a third.
The fifteen most used models
| # | Model | Tokens | Share | CO2e |
|---|---|---|---|---|
| 1 | DeepSeek V4 FlashDeepSeek | 46.7T | 12.3% | 10.3 kt |
| 2 | HY3Tencent | 34.9T | 9.2% | 7.7 kt |
| 3 | MiMo V2.5Xiaomi | 30T | 7.9% | 6.6 kt |
| 4 | GPT 5.6 LunaOpenAI | 24.9T | 6.6% | 5.5 kt |
| 5 | DeepSeek V4 FlashDeepSeek | 23.8T | 6.3% | 5.3 kt |
| 6 | Nemotron 3 Ultra 550B A55BNVIDIAfree | 15.8T | 4.2% | 3.5 kt |
| 7 | GLM 5.2Z-AI | 15T | 4% | 3.3 kt |
| 8 | DeepSeek V4 ProDeepSeek | 9.8T | 2.6% | 2.2 kt |
| 9 | GLM 5.3 FlashZ-AI | 8.1T | 2.1% | 1.8 kt |
| 10 | Claude Opus 5Anthropic | 7.3T | 1.9% | 1.6 kt |
| 11 | Laguna S 2.1Poolsidefree | 7.1T | 1.9% | 1.6 kt |
| 12 | MiniMax M3MiniMax | 7.1T | 1.9% | 1.6 kt |
| 13 | Gemini 3.7 FlashGoogle | 6.5T | 1.7% | 1.4 kt |
| 14 | Kimi K3Moonshot AI | 6.3T | 1.7% | 1.4 kt |
| 15 | Gemini 3.6 FlashGoogle | 6.3T | 1.7% | 1.4 kt |
The table has two details worth pausing on. The dataset treats each build of a model separately, so the April and July builds of DeepSeek V4 Flash are separate rows, and the newer build took over within a month of its release. And two of the fifteen are free to call: NVIDIA's Nemotron 3 Ultra and Poolside's Laguna S are on the list because they cost nothing, which says that at this scale price decides what gets used more than quality does.
What it emits
Our method turns tokens into CO2e in three layers: the energy the chips use for the tokens themselves, the energy the facility uses to keep those chips running and cool, and a small allowance for building the hardware and training the model. On medium-tier factors, with the aggregate split we use for any token count we have not metered (three quarters input, one quarter output, nothing cached), a trillion tokens comes to about 221 tonnes of CO2e. Multiply by August 2026's 379 trillion and you get the 84,000 tonnes above.
The uncertainty band deserves as much attention as the figure. Two 20% uncertainties, one on activity and one on the factors, combine to plus or minus 28%, so August 2026 sits somewhere between 60,000 and 107,000 tonnes. The table's per-model column shares the month's figure out in proportion to tokens, because the public dataset says nothing about which of those tokens were input, output or cached, and nothing about the hardware behind each model. A metered account replaces every one of those assumptions with a count, per call, which is the whole point of measuring rather than estimating.
How to cut it: four levers with numbers
None of these levers changes what you build. Each one changes what a token costs, in energy and in money, and those two move together.
Cache what repeats. System prompts, tool definitions and the documents you paste into every call are read again on every call unless the provider serves them from cache. On our factors a cached input token costs a tenth of a fresh one. Serving half of your input from cache cuts a month's CO2e by about 25%; serving four fifths cuts it by about 40%. OpenAI caches repeated prefixes of 1,024 tokens or more on its own; Anthropic needs a cache marker in the request. Both charge a tenth of the input price for a cache read, which is why this is the lever that gets pulled.
Send routine work to a smaller model. Classification, extraction, routing and short replies rarely need a flagship. A small-tier token costs about a third of a medium-tier token on our factors, so moving half of the traffic to a small model cuts the total by about 28%, and moving all of it cuts it by more than half. The OpenRouter table shows the market has already noticed: the most used models are the Flash and Fast variants.
Ask for shorter answers. Output tokens cost more energy than input tokens, about 0.5 joules against 0.3 on the medium tier, because each one is generated in turn. Bringing output from a quarter of tokens down to 15%, with tighter instructions or a lower maximum length, saves about 5%, which is a small saving and a free one.
Mind the grid. The same tokens on a data centre with a PUE of 1.09 on a 0.345 kilogram grid emit 22% less than on one at 1.2 and 0.42. Region choice is a lever for anyone on a cloud provider that offers it, and it works on every token regardless of the other three.
Pull the first three together, half the input cached, half the traffic on a small model, output at 15%, and a month's CO2e roughly halves.
1B
A billion tokens a month is a mid-sized team on the API. OpenRouter routes a few hundred trillion.
0.3 J per input token, 0.5 J per output token, 0.03 J when cached.
0%
System prompts, documents and tool definitions that repeat between calls.
25%
Output tokens cost more energy than input tokens. Shorter answers move this down.
0%
Classification, extraction and short replies rarely need the flagship.
CO2e a month, after the levers
221 kg
from 221 kg before, plus or minus 28.3 percent on both
The calculator leaves out two things. One is measurement itself: the before figure is an estimate on the aggregate split, and a metered account replaces it with a count, which changes the number and, more often than people expect, changes which lever matters most. The other is what to do with the emissions that remain. Retiring verified credits against the measured figure funds real removals, and on our platform that is a separate step which never touches the measured number, so the report stays honest and the climate action sits on its own line.
What this means for a team
If you run AI through an API, your traffic has the same shape as this curve, only smaller. The share you send to each model changes month by month, the versions change under you, and the emissions follow the tokens. A number you can defend needs three things: the count per call, the factor per model tier, and the method written down. The count is the one part you cannot reconstruct after the fact, so the order matters: meter first, then pull the levers, then retire credits against what is left.
Sources
- 01OpenRouter, LLM Rankings
The public rankings page behind the dataset: tokens per model, updated continuously.
- 02OpenRouter API, rankings-daily dataset
Tokens per model per day for the trailing 360 days; the monthly figures in this post are sums of it, refreshed 8 September 2026. The first month in the window is cut short and is left out.
- 03CrbonFree methodology v1.2
The tier factors, the three layers and the uncertainty band used for every figure on this page.
- 04CrbonFree, OpenRouter case study
The same eleven months with the method worked through in full, month by month.
- 05Anthropic, Prompt caching
How cached prefixes are reused across calls; cache reads are billed at a tenth of the input price.
- 06OpenAI, Prompt caching
Automatic caching of repeated prompt prefixes of 1,024 tokens or more.
Figures attributed to CrbonFree come from methodology v1.2, published in full with every factor and formula. Read the methodology.
About the author
Cory Bergh
Cory leads Crbon Labs, which originates its own climate projects and builds CrbonFree, the platform that measures the footprint of AI usage per token. He was previously VP of Technology and Innovation in the energy industry and holds a BComm, an MBA and the CFA.

