ResearchOpenRoutermarket shareDeepSeek
Stacked bands of colour rising like a tide across the frame, each band a different tone

LLM market share in tokens: a year of OpenRouter

Which models run the most AI traffic, measured in tokens: 1.3 quadrillion tokens through OpenRouter in eleven months, who gained share, and what it emits.

CBCory Bergh, CEO and co-founder, Crbon LabsPublished 9 September 202610 min read

In August 2026, OpenRouter routed 379 trillion tokens, 17x what it routed in October 2025. DeepSeek carried 22.4% of them, OpenAI 10.5%, Tencent 10.3%, Xiaomi 8.5% and Google 7.4%. On our method the month came to about 84,000 tonnes of CO2e, give or take 28%. This post has the curve, the shares month by month, the fifteen most used models, and a calculator for cutting the number.

Market share for language models is usually reported from surveys, which ask companies which vendor they use. That tells you who bought what and very little about how much work gets done, because a company running a billion tokens a day and one running a thousand count the same. Tokens are the unit the models are actually paid in, and OpenRouter publishes them, per model, every day.

Why tokens, and why OpenRouter

OpenRouter is a gateway. A developer sends a request to one endpoint and it goes to whichever provider they chose, so OpenRouter can see which models get used, and it publishes the counts on its rankings page and through a dataset behind it. We sum that dataset by month and by provider, and we run every month through the same method the product uses to turn tokens into CO2e. The numbers below are the sums as of September 8, 2026.

It is a sample rather than the whole market. Traffic that goes straight to OpenAI's or Anthropic's own API never passes through OpenRouter, and neither does anything inside ChatGPT, Gemini or Copilot. The developers who route through a gateway are more price-sensitive than average, and some of the largest applications on it are chat and role-play products. That skews the shares towards cheap and free models. Even so, it is the largest public dataset that counts tokens per model, and its direction of travel is hard to argue with.

The year in one curve

OctDecFebAprJunAug

Monthly tokens on OpenRouter, from the product's public endpoint. Fetched 13 Sep 2026, 17:19 UTC.

Monthly tokens went from 22.3 trillion in October 2025 to 379.4 trillion in August 2026. That is 17x in ten steps, a doubling roughly every 2.5 months, or about 33% growth a month. The 11 complete months shown add up to 1.28 quadrillion tokens. On our factors that is 282,000 tonnes of CO2e, which is what about 61,000 cars emit in a year, from one gateway.

August 2026 alone was 84,000 tonnes, with a band of 60,000 to 107,000, and about 174 gigawatt-hours of electricity at the data centre meter once host power, cluster utilisation and cooling are counted.

Who gained share, and who lost it

Oct: DeepSeek 6.6% (1.5T tokens)Oct: OpenAI 10.1% (2.3T tokens)Oct: Tencent 0% (0 tokens)Oct: Xiaomi 0% (0 tokens)Oct: Google 18.6% (4.2T tokens)Oct: Z-AI 2.5% (559B tokens)Oct: Other 62.2% (13.9T tokens)OctNov: DeepSeek 5% (1.3T tokens)Nov: OpenAI 7.5% (2T tokens)Nov: Tencent 0% (0 tokens)Nov: Xiaomi 0% (0 tokens)Nov: Google 18.5% (4.9T tokens)Nov: Z-AI 2.4% (636B tokens)Nov: Other 66.6% (17.7T tokens)NovDec: DeepSeek 7.5% (1.9T tokens)Dec: OpenAI 11.2% (2.9T tokens)Dec: Tencent 0% (0 tokens)Dec: Xiaomi 2.8% (719B tokens)Dec: Google 22% (5.7T tokens)Dec: Z-AI 1.9% (483B tokens)Dec: Other 54.7% (14.2T tokens)DecJan: DeepSeek 8.1% (2.6T tokens)Jan: OpenAI 11% (3.5T tokens)Jan: Tencent 0% (0 tokens)Jan: Xiaomi 5.5% (1.8T tokens)Jan: Google 23.7% (7.5T tokens)Jan: Z-AI 2.5% (806B tokens)Jan: Other 49.2% (15.6T tokens)JanFeb: DeepSeek 7.1% (3.5T tokens)Feb: OpenAI 10% (5T tokens)Feb: Tencent 0% (0 tokens)Feb: Xiaomi 0.5% (257B tokens)Feb: Google 18.1% (9T tokens)Feb: Z-AI 5.7% (2.8T tokens)Feb: Other 58.7% (29.2T tokens)FebMar: DeepSeek 6.3% (5.3T tokens)Mar: OpenAI 8.6% (7.2T tokens)Mar: Tencent 0% (0 tokens)Mar: Xiaomi 10.5% (8.8T tokens)Mar: Google 14.7% (12.3T tokens)Mar: Z-AI 4.8% (4.1T tokens)Mar: Other 55% (46.1T tokens)MarApr: DeepSeek 5.9% (5.7T tokens)Apr: OpenAI 8.2% (8T tokens)Apr: Tencent 1.9% (1.8T tokens)Apr: Xiaomi 5.1% (4.9T tokens)Apr: Google 14.2% (13.8T tokens)Apr: Z-AI 4.5% (4.4T tokens)Apr: Other 60.2% (58.4T tokens)AprMay: DeepSeek 15.1% (18.6T tokens)May: OpenAI 7.4% (9T tokens)May: Tencent 11.2% (13.8T tokens)May: Xiaomi 2.9% (3.6T tokens)May: Google 13% (16T tokens)May: Z-AI 2.7% (3.3T tokens)May: Other 47.6% (58.5T tokens)MayJun: DeepSeek 17.1% (32.2T tokens)Jun: OpenAI 6.2% (11.6T tokens)Jun: Tencent 8.1% (15.3T tokens)Jun: Xiaomi 9.4% (17.8T tokens)Jun: Google 9.1% (17.2T tokens)Jun: Z-AI 3.7% (6.9T tokens)Jun: Other 46.4% (87.4T tokens)JunJul: DeepSeek 16.6% (41.3T tokens)Jul: OpenAI 5.3% (13.2T tokens)Jul: Tencent 11.9% (29.5T tokens)Jul: Xiaomi 14.8% (36.7T tokens)Jul: Google 7.1% (17.6T tokens)Jul: Z-AI 6.2% (15.5T tokens)Jul: Other 38.1% (94.6T tokens)JulAug: DeepSeek 22.4% (84.9T tokens)Aug: OpenAI 10.5% (40T tokens)Aug: Tencent 10.3% (38.9T tokens)Aug: Xiaomi 8.5% (32.1T tokens)Aug: Google 7.4% (27.9T tokens)Aug: Z-AI 6.7% (25.2T tokens)Aug: Other 34.3% (130T tokens)Aug
  • DeepSeek22.4%
  • OpenAI10.5%
  • Tencent10.3%
  • Xiaomi8.5%
  • Google7.4%
  • Z-AI6.7%
  • Other34.3%
Share of OpenRouter tokens by provider, month by month. The six providers are the largest in the latest complete month; everyone else is in grey. Hover a segment for the month's figure. Fetched 13 Sep 2026, 17:19 UTC.

Google's share shrank. It was the largest named provider through the winter, peaking at 23.7% in January 2026 on Gemini 2.5 Flash, and it ended August 2026 at 7.4%. Its volume still grew, because everything grew, but its slice fell to a third of what it was.

DeepSeek's growth came in one step. Its share sat between 5% and 8% for 6 months, then V4 Flash arrived in April 2026 and the share went to 15% in May, 17% in June and 22.4% in August. Two builds of that one model now carry 18.6% of all traffic on the gateway.

Two providers that were not on the board in March 2026 are on it now. Tencent's HY3 went from nothing in April 2026 to 11% in May and has held around 10% since. Xiaomi's MiMo V2.5 climbed to 14.8% in July 2026 before settling at 8.5%. Together with Z-AI's GLM models, the Chinese labs held 48% of August 2026's tokens.

OpenAI held close to 10% all year, which means its volume rose 17x along with the market. GPT 5.6 Luna is its most used model at 6.6%. Anthropic is the counter-example: Claude 4.5 Sonnet was the second most used model on the gateway in October 2025 at 10%, and Claude Opus 5 sits at 1.9% now. That is a statement about who routes Claude through OpenRouter, not about Claude's use in general.

The other change is concentration. The six largest providers carried 38% of tokens in October 2025 and 66% in August 2026. The "other" bucket, which was 62% of the gateway a year ago, is now a third.

The fifteen most used models

#ModelTokensShareCO2e
1DeepSeek V4 FlashDeepSeek46.7T12.3%10.3 kt
2HY3Tencent34.9T9.2%7.7 kt
3MiMo V2.5Xiaomi30T7.9%6.6 kt
4GPT 5.6 LunaOpenAI24.9T6.6%5.5 kt
5DeepSeek V4 FlashDeepSeek23.8T6.3%5.3 kt
6Nemotron 3 Ultra 550B A55BNVIDIAfree15.8T4.2%3.5 kt
7GLM 5.2Z-AI15T4%3.3 kt
8DeepSeek V4 ProDeepSeek9.8T2.6%2.2 kt
9GLM 5.3 FlashZ-AI8.1T2.1%1.8 kt
10Claude Opus 5Anthropic7.3T1.9%1.6 kt
11Laguna S 2.1Poolsidefree7.1T1.9%1.6 kt
12MiniMax M3MiniMax7.1T1.9%1.6 kt
13Gemini 3.7 FlashGoogle6.5T1.7%1.4 kt
14Kimi K3Moonshot AI6.3T1.7%1.4 kt
15Gemini 3.6 FlashGoogle6.3T1.7%1.4 kt
The fifteen most used models on OpenRouter in Aug 2026, 66.0 percent of the month's tokens between them. CO2e is the month's aggregate figure shared out in proportion, not a per-model measurement.

The table has two details worth pausing on. The dataset treats each build of a model separately, so the April and July builds of DeepSeek V4 Flash are separate rows, and the newer build took over within a month of its release. And two of the fifteen are free to call: NVIDIA's Nemotron 3 Ultra and Poolside's Laguna S are on the list because they cost nothing, which says that at this scale price decides what gets used more than quality does.

What it emits

Our method turns tokens into CO2e in three layers: the energy the chips use for the tokens themselves, the energy the facility uses to keep those chips running and cool, and a small allowance for building the hardware and training the model. On medium-tier factors, with the aggregate split we use for any token count we have not metered (three quarters input, one quarter output, nothing cached), a trillion tokens comes to about 221 tonnes of CO2e. Multiply by August 2026's 379 trillion and you get the 84,000 tonnes above.

The uncertainty band deserves as much attention as the figure. Two 20% uncertainties, one on activity and one on the factors, combine to plus or minus 28%, so August 2026 sits somewhere between 60,000 and 107,000 tonnes. The table's per-model column shares the month's figure out in proportion to tokens, because the public dataset says nothing about which of those tokens were input, output or cached, and nothing about the hardware behind each model. A metered account replaces every one of those assumptions with a count, per call, which is the whole point of measuring rather than estimating.

How to cut it: four levers with numbers

None of these levers changes what you build. Each one changes what a token costs, in energy and in money, and those two move together.

Cache what repeats. System prompts, tool definitions and the documents you paste into every call are read again on every call unless the provider serves them from cache. On our factors a cached input token costs a tenth of a fresh one. Serving half of your input from cache cuts a month's CO2e by about 25%; serving four fifths cuts it by about 40%. OpenAI caches repeated prefixes of 1,024 tokens or more on its own; Anthropic needs a cache marker in the request. Both charge a tenth of the input price for a cache read, which is why this is the lever that gets pulled.

Send routine work to a smaller model. Classification, extraction, routing and short replies rarely need a flagship. A small-tier token costs about a third of a medium-tier token on our factors, so moving half of the traffic to a small model cuts the total by about 28%, and moving all of it cuts it by more than half. The OpenRouter table shows the market has already noticed: the most used models are the Flash and Fast variants.

Ask for shorter answers. Output tokens cost more energy than input tokens, about 0.5 joules against 0.3 on the medium tier, because each one is generated in turn. Bringing output from a quarter of tokens down to 15%, with tighter instructions or a lower maximum length, saves about 5%, which is a small saving and a free one.

Mind the grid. The same tokens on a data centre with a PUE of 1.09 on a 0.345 kilogram grid emit 22% less than on one at 1.2 and 0.42. Region choice is a lever for anyone on a cloud provider that offers it, and it works on every token regardless of the other three.

Pull the first three together, half the input cached, half the traffic on a small model, output at 15%, and a month's CO2e roughly halves.

1B

1M1T

A billion tokens a month is a mid-sized team on the API. OpenRouter routes a few hundred trillion.

0.3 J per input token, 0.5 J per output token, 0.03 J when cached.

0%

0%90%

System prompts, documents and tool definitions that repeat between calls.

25%

5%50%

Output tokens cost more energy than input tokens. Shorter answers move this down.

0%

0%100%

Classification, extraction and short replies rarely need the flagship.

CO2e a month, after the levers

221 kg

from 221 kg before, plus or minus 28.3 percent on both

Before, on the aggregate split221 kg
After221 kg
Saved0.0 kg (0%)
Electricity at the meter, after459 kWh
Tokens1B
CO2e = (uncached input × prefill J + cached input × cached J + output × decode J) ÷ 3,600,000 × 1.18 ÷ 0.30 × PUE × grid, plus 0.028 kg per million tokens for hardware and training.
Methodology v1.2 tier factors, the same ones the product and the case study use. The before figure applies the aggregate split, 75% uncached input and 25% output, which is how the site treats any token count it has not metered. Measured usage replaces every assumption here with a count.

The calculator leaves out two things. One is measurement itself: the before figure is an estimate on the aggregate split, and a metered account replaces it with a count, which changes the number and, more often than people expect, changes which lever matters most. The other is what to do with the emissions that remain. Retiring verified credits against the measured figure funds real removals, and on our platform that is a separate step which never touches the measured number, so the report stays honest and the climate action sits on its own line.

What this means for a team

If you run AI through an API, your traffic has the same shape as this curve, only smaller. The share you send to each model changes month by month, the versions change under you, and the emissions follow the tokens. A number you can defend needs three things: the count per call, the factor per model tier, and the method written down. The count is the one part you cannot reconstruct after the fact, so the order matters: meter first, then pull the levers, then retire credits against what is left.

Sources

  1. 01
    OpenRouter, LLM Rankings

    The public rankings page behind the dataset: tokens per model, updated continuously.

  2. 02
    OpenRouter API, rankings-daily dataset

    Tokens per model per day for the trailing 360 days; the monthly figures in this post are sums of it, refreshed 8 September 2026. The first month in the window is cut short and is left out.

  3. 03
    CrbonFree methodology v1.2

    The tier factors, the three layers and the uncertainty band used for every figure on this page.

  4. 04
    CrbonFree, OpenRouter case study

    The same eleven months with the method worked through in full, month by month.

  5. 05
    Anthropic, Prompt caching

    How cached prefixes are reused across calls; cache reads are billed at a tenth of the input price.

  6. 06
    OpenAI, Prompt caching

    Automatic caching of repeated prompt prefixes of 1,024 tokens or more.

Figures attributed to CrbonFree come from methodology v1.2, published in full with every factor and formula. Read the methodology.

About the author

Cory Bergh

Cory leads Crbon Labs, which originates its own climate projects and builds CrbonFree, the platform that measures the footprint of AI usage per token. He was previously VP of Technology and Innovation in the energy industry and holds a BComm, an MBA and the CFA.

Questions

Questions this post gets asked.

Short answers. The sources above and the methodology have the arithmetic.

Read the methodology
  • By tokens routed through OpenRouter in August 2026, DeepSeek leads with 22.4%, and its V4 Flash model is the single most used model at 18.6% across two versions. OpenAI and Tencent follow at about 10% each. Surveys of businesses tell a different story, because they count companies rather than work done, and OpenRouter is one gateway rather than the whole market.

  • On OpenRouter, DeepSeek V4 Flash. Its July 2026 build alone carried 12.3% of August 2026’s tokens, and the April build another 6.3%. Tencent’s HY3 and Xiaomi’s MiMo V2.5 come next. A year earlier the top model was xAI’s Grok Code Fast 1 at 24%, which shows how quickly the ranking turns over.

  • On OpenRouter, OpenAI’s share of tokens held around 10% for the whole year while the total grew 17x, so its volume grew with the market even as cheaper models took a larger slice. This gateway sees API traffic from developers, not ChatGPT’s consumer use, so it says nothing about the chat product itself.

  • About 379 trillion in August 2026, up from 22 trillion in October 2025. The figure is the monthly sum of OpenRouter’s public rankings dataset, which lists tokens per model per day. Over the 11 complete months in the dataset’s window the total is 1.28 quadrillion tokens.

  • About 221 tonnes of CO2e on our medium-tier factors, with an uncertainty band of plus or minus 28%, when the tokens are three quarters input and one quarter output with nothing cached. Metered traffic replaces that split with a real count, and the answer moves with the model tier, the cache rate and the grid the data centre runs on.

  • Yes, and by more than most levers. A cached input token costs about a tenth of the energy of one read fresh on our factors, so serving half of the input from cache cuts a month’s CO2e by about a quarter, and serving four fifths from cache cuts it by about 40%. Caching also lowers the bill, which is why it tends to get done.

See your own number instead of ours.

However your team already runs AI, that is how you connect. Paste a read-only provider key, drop in an SDK, add the MCP server, install the CLI or switch on the browser extension, and the first measured tokens reach your dashboard within seconds.

Create your free account

Free plan · first tokens in seconds