Google publishes the energy of a single Gemini prompt, 0.24 watt-hours, and companies quote it, but Google does not publish how many prompts it serves, so nobody outside Google can check the figure. Google’s count of the tokens processed by its models was over 3.2 quadrillion in May 2026, and working back from that count, the 0.24 watt-hours holds if a typical prompt carries 1,100 to 2,700 tokens, a plausible size for a chat session, so Google’s math is probably reasonable. This post runs the token count through Crbon Labs methodology v1.2, shows what the boundary and cached tokens do to the total, and explains why a token count can be scaled to a company’s own AI use when a per-prompt figure cannot.
Most companies that sell AI publish one kind of number: the tokens their models process or, less often, the energy of a single request. Google publishes both. It gives a token count every few months, and in August 2025 it published an energy figure per prompt, measured in its own data centres, with the method behind it. We found no other company that publishes both under a published method. With the two side by side, Google’s own disclosures and a published factor set are enough to check one number against the other.
Google’s tokens, month by month
Google calls tokens the fundamental units of data its models process; in English a token is about four characters of text. It counts them across its “surfaces”, the products and APIs that run its models, without breaking the count down by product or model.
| Month | Tokens a month | Where Google said it |
|---|---|---|
| May 2024 | 9.7 trillion | I/O keynote, 20 May 2025, as the figure for a year earlier |
| May 2025 | over 480 trillion | I/O keynote, 20 May 2025 |
| July 2025 | over 980 trillion | Second-quarter earnings call, 23 July 2025 |
| Summer 2025 | 1.3 quadrillion | Gemini at Work, 9 October 2025 |
| May 2026 | over 3.2 quadrillion | I/O keynote, 19 May 2026 |
- All Google surfaces
Crbon Labs Inc. (2026)
From May 2024 to May 2026 the monthly count grew about 330 times, a doubling roughly every three months. Google put the last year at 7 times; on the figures as stated it is 6.7 times. The May 2026 figure is “over” 3.2 quadrillion, so 330 times is a floor.
Google also gives a second, narrower count on its earnings calls: tokens a minute through its model APIs, the part developers and companies buy directly. It was 7 billion a minute in October 2025, over 10 billion in February 2026, more than 16 billion in April 2026 and approximately 22 billion in July 2026. At 22 billion a minute, a month of API use is about 960 trillion tokens. On the one day Google gave the monthly total and the API rate together, 19 May 2026, its API rate of roughly 19 billion a minute came to about a quarter of the monthly total.
How much energy a Gemini prompt uses
In August 2025 Google published a paper on the energy, carbon and water of serving Gemini, measured across its production fleet. For the median Gemini Apps text prompt in May 2025 it found 0.24 watt-hours of energy, 0.03 grams of CO2e and 0.26 millilitres of water. The energy splits four ways: the active accelerators drew 58%, the host machine’s CPU and memory 25%, idle machines held ready for traffic 10%, and data centre overhead 8%, at a fleet power usage effectiveness of 1.09, meaning the building draws 9% more power than its computers.
The paper also reports what a narrower method, which it calls the existing approach, gives for the same prompt. Count only the active chips, and only for prompts served in the 10% most efficient data centres, and the median prompt comes to 0.10 watt-hours. Counting the host, idle machines and overhead in those same data centres brings it to 0.17 watt-hours, and averaging across the whole fleet brings it to 0.24 watt-hours. The comprehensive figure is 2.4 times the narrow one: 1.7 times for what is counted and 1.4 times for which data centres are sampled.
- Active accelerators
- Host CPU and memory
- Idle machines
- Data centre overhead
Crbon Labs Inc. (2026)
The carbon figure is market-based. Google applies its 2024 fleet factor after clean energy purchases, 94 grams per kWh, and adds the embodied emissions of its hardware. On the grids themselves the same electricity carried 345 grams per kWh in 2024, the location-based factor, and at that rate 0.24 watt-hours comes to about 0.08 grams, or 0.09 with the hardware. The paper leaves out model training. It does not say how many tokens the median prompt contains, or how many prompts Gemini serves.
Energy per median prompt fell 33 times in the twelve months to May 2025, from 7.92 watt-hours in May 2024, which Google credits mostly to more efficient models and partly to better use of its machines. By Google’s own note, the figures have not been verified by an independent third party. Our post on how much energy AI uses sets Google’s figure beside the per-query figures other providers have published.
How much energy Google’s AI uses in a month
Crbon Labs methodology v1.2, which we use throughout, converts tokens to energy and carbon in three layers. The first layer is the electricity the chips draw while they work: a joule figure per token for input, output and cached reads, multiplied by the data centre overhead and the grid factor. The second adds the power of the host machine around the chips and the idle capacity of a cluster running at about 30% utilisation. The third layer adds the manufacture of the hardware and the training of the model, spread over the tokens they serve. Each factor is listed on the methodology page.
For Google’s fleet the methodology uses a power usage effectiveness of 1.09 and 345 grams of CO2e per kWh, Google’s own location-based factor for 2024. Google does not say which models served its tokens, so we run the month on two tiers: Gemini Flash, at 0.08 joules per input token and 0.5 joules per output token, and Gemini Pro with reasoning, at 0.35 joules per input token and 2.5 joules per output token. Both use the split the methodology applies to an aggregate count, 75% input and 25% output, with nothing cached.
| May 2026, 3.2 quadrillion tokens | Gemini Flash tier | Gemini Pro reasoning tier |
|---|---|---|
| Active energy of the chips | 164 GWh | 789 GWh |
| Energy at the data centre | 705 GWh | 3,382 GWh |
| Layer 1, tonnes of CO2e | 61,839 | 296,662 |
| Layer 2, tonnes of CO2e | 243,235 | 1,166,869 |
| Layer 3, tonnes of CO2e | 332,835 | 1,256,469 |
| Uncertainty band, plus or minus 28.3% | 238,709 to 426,960 | 901,140 to 1,611,799 |
- Active compute
- Host power and idle capacity
- Embodied and training
Crbon Labs Inc. (2026)
Google’s environmental report helps decide which tier fits most of the month. Its data centres used 42.4 TWh in 2025, about 3,535 GWh in an average month, for everything they run, including Search, cloud customers’ workloads and model training. On the Flash tier, May 2026’s tokens need 705 GWh at the data centre, about a fifth of that. On the Pro reasoning tier they would need 3,382 GWh, almost all of it. On average, then, Google’s tokens must cost well under the Pro reasoning tier, and the Flash tier is the closer description, which is why we lead with it. The comparison sets a May 2026 rate against 2025 electricity, which itself grew 38% that year, so it gives a sense of scale.
Per day, the Flash tier puts May 2026 at about 23 GWh of data centre electricity and 10,900 tonnes of CO2e, or 965 MWh and 456 tonnes an hour. For May 2025, the month Google measured its median prompt, the same tier gives 105.8 GWh and 49,925 tonnes for 480 trillion tokens.
Where Google’s per-prompt figure meets its token count
To turn Google’s per-prompt figure into a monthly total you need a count of prompts, and Google does not disclose one. It reports users, 950 million a month for the Gemini app in July 2026, and growth in requests, and its paper divides by a prompt count it keeps to itself. So the check has to run the other way round: our token-based total tells us how many tokens a median prompt would have to contain for Google’s 0.24 watt-hours to be right.
Google’s 0.24 watt-hours covers the same ground as our second layer: the chips, the host, idle machines and overhead. Its 0.10 watt-hours covers only the active chips, like our active energy.
| Tier | Tokens per prompt for our data centre energy to equal 0.24 Wh | Tokens per prompt for our active energy to equal 0.10 Wh |
|---|---|---|
| Gemini Flash | 1,089 | 1,946 |
| Gemini Pro reasoning | 227 | 406 |
Google has not published the real size of a median prompt. Watershed’s open framework for measuring emissions from AI usage, which our methodology adopts, derives its own defaults from Google’s 0.24 watt-hours and an assumed prompt of about 500 tokens.
Here is what the table says in words. On the Flash tier, Google’s 0.24 watt-hours and our data centre figure agree if a median prompt is about 1,100 tokens, and its 0.10 watt-hours and our chip-only figure agree at about 1,900 tokens. Watershed assumes 500 tokens. The two prompt sizes differ because the step from the chips to the whole data centre is where our defaults and Google’s measurements part company: Google measured that step at 1.7 times on its fleet, while our second layer applies 4.3 times, because it assumes a cluster runs at 30% utilisation, the default for providers that publish no measurement of their own. Google’s idle machines take about 10% of its energy, not most of it. With Google’s measured ratios in place of our defaults, the Flash-tier month falls to 282 gigawatt-hours at the data centre and 186,776 tonnes of CO2e, and the prompt size at which its 0.24 watt-hours agrees with its token count rises to about 2,700 tokens.
That is the answer to the question in the title. Google’s per-prompt figure and its token count are consistent with each other if a median Gemini prompt carries somewhere between 1,100 and 2,700 tokens, which is plausible for a chat session, where each message carries the conversation’s history and instructions, and two to five times the 500 tokens the open framework assumes. Nobody outside Google can narrow that range, because the one number that would settle it, the count of prompts, is the one Google does not publish. Until it does, 0.24 watt-hours is a figure to quote, not one to scale.
What the boundary does to the answer
Google’s two figures differ by 2.4 times. The gap between the Gemini Flash and Gemini Pro tiers in our methodology is larger, 4.8 times in active energy. A report makes more than one boundary choice, though, and they stack. Counted four ways, the same May 2026 month on the Flash tier runs from about 17,000 to 333,000 tonnes.
| What is counted | Tonnes of CO2e |
|---|---|
| Layer 1 on Google’s market-based factor, 94 g per kWh | 16,849 |
| Layer 1 on the location-based factor, 345 g per kWh | 61,839 |
| Layer 2, adding host power and idle capacity | 243,235 |
| Layer 3, adding hardware and training | 332,835 |
From the first row to the last is 19.8 times, on the same tokens and the same model tier. Switching from the Flash tier to the Pro reasoning tier changes it 4.8 times. The number depends more on the boundary than on the model, and a figure for AI emissions published without its boundary cannot be compared with any other. Google’s paper makes the same argument for standard boundaries that count everything a prompt uses.
What cached tokens change
A cached token is part of a prompt the model has already processed and kept in memory, typically the instructions or documents that open a conversation and repeat on every turn. Reading it again skips most of the computation. Google added implicit caching to Gemini 2.5 models in May 2025, and its Vertex AI bills cached tokens at 10% of the standard input price on Gemini 2.5 and later models. Our methodology charges a cached read a tenth of the energy of an uncached input token, which happens to match the ratio in Google’s price.
Google publishes no cache hit rate, so we show two scenarios on the May 2026 month. If half the input tokens were read from cache, the Flash-tier month falls from 332,835 to 297,336 tonnes, 10.7% less. At 70% it falls to 283,136 tonnes, 14.9% less, about 50,000 tonnes. The effect is modest on the Gemini tiers because most of their energy goes into writing output, which caching does not touch. On our medium tier, where input costs more, the same 70% cuts the month by 35%.
Watershed’s framework notes that cached input is cheaper and lets a provider disclose its own factor for it. Its default equation has no cached term, so on its defaults a month served 70% from cache is reported at the uncached figure. Methodology v1.2 carries a default for cached reads, and Google’s API reports the cached count on every request (cachedContentTokenCount), so a company can apply it to its own traffic.
Who else publishes energy per prompt and tokens served
OpenAI publishes an energy figure per query and a token rate, both in a looser form than Google’s. Its chief executive put the average ChatGPT query at about 0.34 watt-hours in June 2025, without a method or a boundary, and the company said in March 2026 that its APIs process more than 15 billion tokens a minute. Microsoft gives token counts on its earnings calls, and its July 2026 sustainability report cites its researchers’ finding that optimised inference can use less than one watt-hour per query, with no volume beside it. Mistral published a lifecycle analysis in July 2025, 1.14 grams of CO2e for a 400-token response, with no token volume. We found neither figure from Anthropic, and only a leaked token count from Meta. Watershed’s framework puts it this way: “Google is the only frontier-model provider to publish per-prompt figures at this granularity.”
What Google’s two numbers leave a buyer to do
What you have read is one company’s two disclosures set against each other. Google publishes an energy figure per prompt, 0.24 watt-hours, and a token count, over 3.2 quadrillion in May 2026, and through our methodology the tokens come to about 705 gigawatt-hours and 333,000 tonnes of CO2e on the Gemini Flash tier for that month. The two disclosures agree only for a median prompt of 1,100 to 2,700 tokens, and Google’s prompt count, which would settle it, is not published.
What it means for anyone reporting their own AI use is that a per-prompt figure from a provider cannot be scaled to a company’s usage, because nobody outside the provider knows what a prompt weighs. Tokens can be scaled, because every API returns them. The number to keep is tokens per model, with the boundary of the factor set that converts them named beside the result: chips only, the whole data centre, or the full lifecycle, and location-based or market-based for the grid.
The electricity behind the tokens Google serves for its own products is Google’s Scope 2, and the hardware is its Scope 3. Tokens a company buys through the Gemini API or Vertex AI, about a quarter of the total on the May 2026 figures, are a purchased service and belong in the buyer’s Scope 3, Category 1, as our guide to Scope 3 emissions from AI sets out. A buyer reporting on the location-based method uses the grid factor, which on Google’s own 2024 figures is 3.7 times the market-based one it applies to prompts; our post on market-based and location-based emissions explains the two methods. The token counts come back with every Gemini response, per request and per model, and they are what CrbonFree meters for Vertex AI and the other providers it connects.
Sources
- 01Elsworth et al., “Measuring the environmental impact of delivering AI at Google Scale”, arXiv, 21 August 2025
The median Gemini Apps text prompt in May 2025: 0.24 Wh, 0.03 g CO2e market-based with embodied hardware, 0.26 mL; the 58, 25, 10 and 8% split; 0.10 Wh on the narrow approach; the 2024 location-based factor of 345 g per kWh; no token count per prompt.
- 02Amin Vahdat and Jeff Dean, “How much energy does Google’s AI use? We did the math”, Google Cloud blog, 21 August 2025
The same figures for a general reader, and Google’s note that they have not been verified by an independent third party.
- 03Sundar Pichai, “Google I/O 2025: From research to reality”, 20 May 2025
9.7 trillion tokens a month a year earlier; over 480 trillion in May 2025.
- 04Sundar Pichai, “Q2 earnings call: CEO’s remarks”, 23 July 2025
Over 980 trillion monthly tokens across Google’s surfaces.
- 05Sundar Pichai, remarks at Gemini at Work, 9 October 2025
1.3 quadrillion monthly tokens, reached in the summer of 2025.
- 06Sundar Pichai, “I/O 2026: Welcome to the agentic Gemini era”, 19 May 2026
Over 3.2 quadrillion tokens a month across Google’s surfaces; model APIs at roughly 19 billion tokens a minute.
- 07Sundar Pichai, “Q2 2026 earnings call: Remarks from our CEO”, 22 July 2026
Model APIs at approximately 22 billion tokens a minute, up from 16 billion a quarter earlier; the Gemini app at 950 million monthly active users.
- 08Google, 2026 Environmental Report
Data centre electricity of 42,415,800 MWh in 2025, fleet PUE of 1.09, and the May 2024 baseline of 7.92 Wh per median prompt.
- 09Bistline et al., “Estimating GHG Emissions from AI Use: Framework for Corporate-Level Measurement”, Watershed, August 2026
The open framework methodology v1.2 adopts. It allows a provider-disclosed factor for cached input and sets no default, and derives its hyperscaler defaults from Google’s 0.24 Wh and an assumed prompt of about 500 tokens.
- 10Google Cloud, “Save costs and decrease latency while using Gemini with Vertex AI context caching”, 15 October 2025
Cached tokens billed at 10% of the standard input price on Gemini 2.5 and later models, with implicit caching applied automatically.
- 11Sam Altman, “The Gentle Singularity”, 10 June 2025
About 0.34 Wh for an average ChatGPT query, with no method or boundary described.
- 12OpenAI, “OpenAI raises $122 billion to accelerate the next phase of AI”, 31 March 2026
OpenAI’s APIs process more than 15 billion tokens a minute.
- 13Mistral AI, “Our contribution to a global environmental standard for AI”, 22 July 2025
A lifecycle analysis of a 400-token Le Chat response: 1.14 g CO2e and 45 mL of water. No token volume.
Figures attributed to CrbonFree come from methodology v1.2, published in full with every factor and formula. Read the methodology.
About the author
Cory Bergh
Cory leads Crbon Labs, which originates its own climate projects and builds CrbonFree, the platform that measures the footprint of AI usage per token. He was previously VP of Technology and Innovation in the energy industry and holds a BComm, an MBA and the CFA.
Share this post
Send it to whoever asked the question.

