ResearchGoogleGeminitokens
A single ring painted on an off-white ground, one half a dense sweep of magenta flecked with gold and orange, the other half a sparse trail of pale blue and green particles, the dense half and the pale half joined into one loop

How much energy does Gemini use? Google’s own math, checked

Google measured 0.24 Wh for a median Gemini prompt in May 2025, and its AI ran over 3.2 quadrillion tokens in May 2026. We checked one number against the other.

CBCory Bergh, CEO and co-founder, Crbon LabsPublished 23 September 202621 min read
Share

Google publishes the energy of a single Gemini prompt, 0.24 watt-hours, and companies quote it, but Google does not publish how many prompts it serves, so nobody outside Google can check the figure. Google’s count of the tokens processed by its models was over 3.2 quadrillion in May 2026, and working back from that count, the 0.24 watt-hours holds if a typical prompt carries 1,100 to 2,700 tokens, a plausible size for a chat session, so Google’s math is probably reasonable. This post runs the token count through Crbon Labs methodology v1.2, shows what the boundary and cached tokens do to the total, and explains why a token count can be scaled to a company’s own AI use when a per-prompt figure cannot.

Most companies that sell AI publish one kind of number: the tokens their models process or, less often, the energy of a single request. Google publishes both. It gives a token count every few months, and in August 2025 it published an energy figure per prompt, measured in its own data centres, with the method behind it. We found no other company that publishes both under a published method. With the two side by side, Google’s own disclosures and a published factor set are enough to check one number against the other.

Google’s tokens, month by month

Google calls tokens the fundamental units of data its models process; in English a token is about four characters of text. It counts them across its “surfaces”, the products and APIs that run its models, without breaking the count down by product or model.

MonthTokens a monthWhere Google said it
May 20249.7 trillionI/O keynote, 20 May 2025, as the figure for a year earlier
May 2025over 480 trillionI/O keynote, 20 May 2025
July 2025over 980 trillionSecond-quarter earnings call, 23 July 2025
Summer 20251.3 quadrillionGemini at Work, 9 October 2025
May 2026over 3.2 quadrillionI/O keynote, 19 May 2026
trillion tokens a monthMay 2024: All Google surfaces 9.7 trillion tokens a month. Stated at I/O on 20 May 2025 as the figure for a year earlier, across products and APIs9.7May 2024May 2025: All Google surfaces 480 trillion tokens a month. Over 480 trillion, I/O, 20 May 2025480May 2025July 2025: All Google surfaces 980 trillion tokens a month. Over 980 trillion, second-quarter earnings call, 23 July 2025980July 2025Summer 2025: All Google surfaces 1,300 trillion tokens a month. Reached in the summer of 2025, announced at Gemini at Work on 9 October 20251,300Summer 2025May 2026: All Google surfaces 3,200 trillion tokens a month. Over 3.2 quadrillion, I/O, 19 May 20263,200May 2026
  • All Google surfaces

Crbon Labs Inc. (2026)

Tokens a month across Google’s products and APIs, in trillions, as Google stated them. The scale is linear, so May 2024’s 9.7 trillion is 0.3% of May 2026’s 3.2 quadrillion. Hover a column for the source.

From May 2024 to May 2026 the monthly count grew about 330 times, a doubling roughly every three months. Google put the last year at 7 times; on the figures as stated it is 6.7 times. The May 2026 figure is “over” 3.2 quadrillion, so 330 times is a floor.

Google also gives a second, narrower count on its earnings calls: tokens a minute through its model APIs, the part developers and companies buy directly. It was 7 billion a minute in October 2025, over 10 billion in February 2026, more than 16 billion in April 2026 and approximately 22 billion in July 2026. At 22 billion a minute, a month of API use is about 960 trillion tokens. On the one day Google gave the monthly total and the API rate together, 19 May 2026, its API rate of roughly 19 billion a minute came to about a quarter of the monthly total.

How much energy a Gemini prompt uses

In August 2025 Google published a paper on the energy, carbon and water of serving Gemini, measured across its production fleet. For the median Gemini Apps text prompt in May 2025 it found 0.24 watt-hours of energy, 0.03 grams of CO2e and 0.26 millilitres of water. The energy splits four ways: the active accelerators drew 58%, the host machine’s CPU and memory 25%, idle machines held ready for traffic 10%, and data centre overhead 8%, at a fleet power usage effectiveness of 1.09, meaning the building draws 9% more power than its computers.

The paper also reports what a narrower method, which it calls the existing approach, gives for the same prompt. Count only the active chips, and only for prompts served in the 10% most efficient data centres, and the median prompt comes to 0.10 watt-hours. Counting the host, idle machines and overhead in those same data centres brings it to 0.17 watt-hours, and averaging across the whole fleet brings it to 0.24 watt-hours. The comprehensive figure is 2.4 times the narrow one: 1.7 times for what is counted and 1.4 times for which data centres are sampled.

Wh per median promptChips only, best sites: Active accelerators 0.1 Wh per median prompt. The narrow approach: active accelerators only, prompts in the 10% most efficient data centres0.1Chips only, best sitesWhole stack, best sites: Active accelerators 0.1 Wh per median prompt. The same efficient data centresWhole stack, best sites: Host CPU and memory 0.04 Wh per median prompt. Measured, left out of the narrow approachWhole stack, best sites: Idle machines 0.02 Wh per median prompt. Measured, left out of the narrow approachWhole stack, best sites: Data centre overhead 0.01 Wh per median prompt. Measured, left out of the narrow approach0.17Whole stack, best sitesWhole stack, fleet: Active accelerators 0.14 Wh per median prompt. 58% of the totalWhole stack, fleet: Host CPU and memory 0.06 Wh per median prompt. 25% of the totalWhole stack, fleet: Idle machines 0.02 Wh per median prompt. 10% of the totalWhole stack, fleet: Data centre overhead 0.02 Wh per median prompt. 8% of the total, at a fleet power usage effectiveness of 1.090.24Whole stack, fleet
  • Active accelerators
  • Host CPU and memory
  • Idle machines
  • Data centre overhead

Crbon Labs Inc. (2026)

Watt-hours for the median Gemini Apps text prompt in May 2025, from Table 1 of Google’s paper. The first column is the narrow approach: active chips only, in the 10% most efficient data centres. The second column adds the host, idle machines and overhead in the same data centres, and the third column averages all four across the fleet. Hover a segment for the figure.

The carbon figure is market-based. Google applies its 2024 fleet factor after clean energy purchases, 94 grams per kWh, and adds the embodied emissions of its hardware. On the grids themselves the same electricity carried 345 grams per kWh in 2024, the location-based factor, and at that rate 0.24 watt-hours comes to about 0.08 grams, or 0.09 with the hardware. The paper leaves out model training. It does not say how many tokens the median prompt contains, or how many prompts Gemini serves.

Energy per median prompt fell 33 times in the twelve months to May 2025, from 7.92 watt-hours in May 2024, which Google credits mostly to more efficient models and partly to better use of its machines. By Google’s own note, the figures have not been verified by an independent third party. Our post on how much energy AI uses sets Google’s figure beside the per-query figures other providers have published.

How much energy Google’s AI uses in a month

Crbon Labs methodology v1.2, which we use throughout, converts tokens to energy and carbon in three layers. The first layer is the electricity the chips draw while they work: a joule figure per token for input, output and cached reads, multiplied by the data centre overhead and the grid factor. The second adds the power of the host machine around the chips and the idle capacity of a cluster running at about 30% utilisation. The third layer adds the manufacture of the hardware and the training of the model, spread over the tokens they serve. Each factor is listed on the methodology page.

For Google’s fleet the methodology uses a power usage effectiveness of 1.09 and 345 grams of CO2e per kWh, Google’s own location-based factor for 2024. Google does not say which models served its tokens, so we run the month on two tiers: Gemini Flash, at 0.08 joules per input token and 0.5 joules per output token, and Gemini Pro with reasoning, at 0.35 joules per input token and 2.5 joules per output token. Both use the split the methodology applies to an aggregate count, 75% input and 25% output, with nothing cached.

May 2026, 3.2 quadrillion tokensGemini Flash tierGemini Pro reasoning tier
Active energy of the chips164 GWh789 GWh
Energy at the data centre705 GWh3,382 GWh
Layer 1, tonnes of CO2e61,839296,662
Layer 2, tonnes of CO2e243,2351,166,869
Layer 3, tonnes of CO2e332,8351,256,469
Uncertainty band, plus or minus 28.3%238,709 to 426,960901,140 to 1,611,799
tonnes of CO2eFlash tier: Active compute 61,839 tonnes of CO2e. 164.4 GWh of active compute at PUE 1.09 and 0.345 kg per kWhFlash tier: Host power and idle capacity 181,396 tonnes of CO2e. The host power factor of 1.18 over 30% utilisation, 705.0 GWh at the data centreFlash tier: Embodied and training 89,600 tonnes of CO2e. 0.028 kg per million tokens for hardware and training332,835Flash tierFlash, 70% cached: Active compute 49,204 tonnes of CO2e. 70% of input tokens read from cache at 0.008 joules eachFlash, 70% cached: Host power and idle capacity 144,332 tonnes of CO2e. 561.0 GWh at the data centreFlash, 70% cached: Embodied and training 89,600 tonnes of CO2e. The adders do not shrink with caching283,136Flash, 70% cachedPro reasoning tier: Active compute 296,662 tonnes of CO2e. 788.9 GWh of active computePro reasoning tier: Host power and idle capacity 870,207 tonnes of CO2e. 3,382.2 GWh at the data centrePro reasoning tier: Embodied and training 89,600 tonnes of CO2e. The same adders per token1,256,469Pro reasoning tier
  • Active compute
  • Host power and idle capacity
  • Embodied and training

Crbon Labs Inc. (2026)

Google’s May 2026 month, over 3.2 quadrillion tokens, in the three layers of methodology v1.2, in tonnes of CO2e: on the Gemini Flash tier, the same with 70% of input tokens read from cache, and on the Gemini Pro reasoning tier. Most of each column is the second layer, the host power and idle capacity around the chips. Hover a segment for the figure.

Google’s environmental report helps decide which tier fits most of the month. Its data centres used 42.4 TWh in 2025, about 3,535 GWh in an average month, for everything they run, including Search, cloud customers’ workloads and model training. On the Flash tier, May 2026’s tokens need 705 GWh at the data centre, about a fifth of that. On the Pro reasoning tier they would need 3,382 GWh, almost all of it. On average, then, Google’s tokens must cost well under the Pro reasoning tier, and the Flash tier is the closer description, which is why we lead with it. The comparison sets a May 2026 rate against 2025 electricity, which itself grew 38% that year, so it gives a sense of scale.

Per day, the Flash tier puts May 2026 at about 23 GWh of data centre electricity and 10,900 tonnes of CO2e, or 965 MWh and 456 tonnes an hour. For May 2025, the month Google measured its median prompt, the same tier gives 105.8 GWh and 49,925 tonnes for 480 trillion tokens.

Where Google’s per-prompt figure meets its token count

To turn Google’s per-prompt figure into a monthly total you need a count of prompts, and Google does not disclose one. It reports users, 950 million a month for the Gemini app in July 2026, and growth in requests, and its paper divides by a prompt count it keeps to itself. So the check has to run the other way round: our token-based total tells us how many tokens a median prompt would have to contain for Google’s 0.24 watt-hours to be right.

Google’s 0.24 watt-hours covers the same ground as our second layer: the chips, the host, idle machines and overhead. Its 0.10 watt-hours covers only the active chips, like our active energy.

TierTokens per prompt for our data centre energy to equal 0.24 WhTokens per prompt for our active energy to equal 0.10 Wh
Gemini Flash1,0891,946
Gemini Pro reasoning227406

Google has not published the real size of a median prompt. Watershed’s open framework for measuring emissions from AI usage, which our methodology adopts, derives its own defaults from Google’s 0.24 watt-hours and an assumed prompt of about 500 tokens.

Here is what the table says in words. On the Flash tier, Google’s 0.24 watt-hours and our data centre figure agree if a median prompt is about 1,100 tokens, and its 0.10 watt-hours and our chip-only figure agree at about 1,900 tokens. Watershed assumes 500 tokens. The two prompt sizes differ because the step from the chips to the whole data centre is where our defaults and Google’s measurements part company: Google measured that step at 1.7 times on its fleet, while our second layer applies 4.3 times, because it assumes a cluster runs at 30% utilisation, the default for providers that publish no measurement of their own. Google’s idle machines take about 10% of its energy, not most of it. With Google’s measured ratios in place of our defaults, the Flash-tier month falls to 282 gigawatt-hours at the data centre and 186,776 tonnes of CO2e, and the prompt size at which its 0.24 watt-hours agrees with its token count rises to about 2,700 tokens.

That is the answer to the question in the title. Google’s per-prompt figure and its token count are consistent with each other if a median Gemini prompt carries somewhere between 1,100 and 2,700 tokens, which is plausible for a chat session, where each message carries the conversation’s history and instructions, and two to five times the 500 tokens the open framework assumes. Nobody outside Google can narrow that range, because the one number that would settle it, the count of prompts, is the one Google does not publish. Until it does, 0.24 watt-hours is a figure to quote, not one to scale.

What the boundary does to the answer

Google’s two figures differ by 2.4 times. The gap between the Gemini Flash and Gemini Pro tiers in our methodology is larger, 4.8 times in active energy. A report makes more than one boundary choice, though, and they stack. Counted four ways, the same May 2026 month on the Flash tier runs from about 17,000 to 333,000 tonnes.

What is countedTonnes of CO2e
Layer 1 on Google’s market-based factor, 94 g per kWh16,849
Layer 1 on the location-based factor, 345 g per kWh61,839
Layer 2, adding host power and idle capacity243,235
Layer 3, adding hardware and training332,835

From the first row to the last is 19.8 times, on the same tokens and the same model tier. Switching from the Flash tier to the Pro reasoning tier changes it 4.8 times. The number depends more on the boundary than on the model, and a figure for AI emissions published without its boundary cannot be compared with any other. Google’s paper makes the same argument for standard boundaries that count everything a prompt uses.

What cached tokens change

A cached token is part of a prompt the model has already processed and kept in memory, typically the instructions or documents that open a conversation and repeat on every turn. Reading it again skips most of the computation. Google added implicit caching to Gemini 2.5 models in May 2025, and its Vertex AI bills cached tokens at 10% of the standard input price on Gemini 2.5 and later models. Our methodology charges a cached read a tenth of the energy of an uncached input token, which happens to match the ratio in Google’s price.

Google publishes no cache hit rate, so we show two scenarios on the May 2026 month. If half the input tokens were read from cache, the Flash-tier month falls from 332,835 to 297,336 tonnes, 10.7% less. At 70% it falls to 283,136 tonnes, 14.9% less, about 50,000 tonnes. The effect is modest on the Gemini tiers because most of their energy goes into writing output, which caching does not touch. On our medium tier, where input costs more, the same 70% cuts the month by 35%.

Watershed’s framework notes that cached input is cheaper and lets a provider disclose its own factor for it. Its default equation has no cached term, so on its defaults a month served 70% from cache is reported at the uncached figure. Methodology v1.2 carries a default for cached reads, and Google’s API reports the cached count on every request (cachedContentTokenCount), so a company can apply it to its own traffic.

Who else publishes energy per prompt and tokens served

OpenAI publishes an energy figure per query and a token rate, both in a looser form than Google’s. Its chief executive put the average ChatGPT query at about 0.34 watt-hours in June 2025, without a method or a boundary, and the company said in March 2026 that its APIs process more than 15 billion tokens a minute. Microsoft gives token counts on its earnings calls, and its July 2026 sustainability report cites its researchers’ finding that optimised inference can use less than one watt-hour per query, with no volume beside it. Mistral published a lifecycle analysis in July 2025, 1.14 grams of CO2e for a 400-token response, with no token volume. We found neither figure from Anthropic, and only a leaked token count from Meta. Watershed’s framework puts it this way: “Google is the only frontier-model provider to publish per-prompt figures at this granularity.”

What Google’s two numbers leave a buyer to do

What you have read is one company’s two disclosures set against each other. Google publishes an energy figure per prompt, 0.24 watt-hours, and a token count, over 3.2 quadrillion in May 2026, and through our methodology the tokens come to about 705 gigawatt-hours and 333,000 tonnes of CO2e on the Gemini Flash tier for that month. The two disclosures agree only for a median prompt of 1,100 to 2,700 tokens, and Google’s prompt count, which would settle it, is not published.

What it means for anyone reporting their own AI use is that a per-prompt figure from a provider cannot be scaled to a company’s usage, because nobody outside the provider knows what a prompt weighs. Tokens can be scaled, because every API returns them. The number to keep is tokens per model, with the boundary of the factor set that converts them named beside the result: chips only, the whole data centre, or the full lifecycle, and location-based or market-based for the grid.

The electricity behind the tokens Google serves for its own products is Google’s Scope 2, and the hardware is its Scope 3. Tokens a company buys through the Gemini API or Vertex AI, about a quarter of the total on the May 2026 figures, are a purchased service and belong in the buyer’s Scope 3, Category 1, as our guide to Scope 3 emissions from AI sets out. A buyer reporting on the location-based method uses the grid factor, which on Google’s own 2024 figures is 3.7 times the market-based one it applies to prompts; our post on market-based and location-based emissions explains the two methods. The token counts come back with every Gemini response, per request and per model, and they are what CrbonFree meters for Vertex AI and the other providers it connects.

Sources

  1. 01
    Elsworth et al., “Measuring the environmental impact of delivering AI at Google Scale”, arXiv, 21 August 2025

    The median Gemini Apps text prompt in May 2025: 0.24 Wh, 0.03 g CO2e market-based with embodied hardware, 0.26 mL; the 58, 25, 10 and 8% split; 0.10 Wh on the narrow approach; the 2024 location-based factor of 345 g per kWh; no token count per prompt.

  2. 02
    Amin Vahdat and Jeff Dean, “How much energy does Google’s AI use? We did the math”, Google Cloud blog, 21 August 2025

    The same figures for a general reader, and Google’s note that they have not been verified by an independent third party.

  3. 03
    Sundar Pichai, “Google I/O 2025: From research to reality”, 20 May 2025

    9.7 trillion tokens a month a year earlier; over 480 trillion in May 2025.

  4. 04
    Sundar Pichai, “Q2 earnings call: CEO’s remarks”, 23 July 2025

    Over 980 trillion monthly tokens across Google’s surfaces.

  5. 05
    Sundar Pichai, remarks at Gemini at Work, 9 October 2025

    1.3 quadrillion monthly tokens, reached in the summer of 2025.

  6. 06
    Sundar Pichai, “I/O 2026: Welcome to the agentic Gemini era”, 19 May 2026

    Over 3.2 quadrillion tokens a month across Google’s surfaces; model APIs at roughly 19 billion tokens a minute.

  7. 07
    Sundar Pichai, “Q2 2026 earnings call: Remarks from our CEO”, 22 July 2026

    Model APIs at approximately 22 billion tokens a minute, up from 16 billion a quarter earlier; the Gemini app at 950 million monthly active users.

  8. 08
    Google, 2026 Environmental Report

    Data centre electricity of 42,415,800 MWh in 2025, fleet PUE of 1.09, and the May 2024 baseline of 7.92 Wh per median prompt.

  9. 09
    Bistline et al., “Estimating GHG Emissions from AI Use: Framework for Corporate-Level Measurement”, Watershed, August 2026

    The open framework methodology v1.2 adopts. It allows a provider-disclosed factor for cached input and sets no default, and derives its hyperscaler defaults from Google’s 0.24 Wh and an assumed prompt of about 500 tokens.

  10. 10
    Google Cloud, “Save costs and decrease latency while using Gemini with Vertex AI context caching”, 15 October 2025

    Cached tokens billed at 10% of the standard input price on Gemini 2.5 and later models, with implicit caching applied automatically.

  11. 11
    Sam Altman, “The Gentle Singularity”, 10 June 2025

    About 0.34 Wh for an average ChatGPT query, with no method or boundary described.

  12. 12
    OpenAI, “OpenAI raises $122 billion to accelerate the next phase of AI”, 31 March 2026

    OpenAI’s APIs process more than 15 billion tokens a minute.

  13. 13
    Mistral AI, “Our contribution to a global environmental standard for AI”, 22 July 2025

    A lifecycle analysis of a 400-token Le Chat response: 1.14 g CO2e and 45 mL of water. No token volume.

Figures attributed to CrbonFree come from methodology v1.2, published in full with every factor and formula. Read the methodology.

About the author

Cory Bergh

Cory leads Crbon Labs, which originates its own climate projects and builds CrbonFree, the platform that measures the footprint of AI usage per token. He was previously VP of Technology and Innovation in the energy industry and holds a BComm, an MBA and the CFA.

Share this post

Send it to whoever asked the question.

Questions

Questions this post gets asked.

Short answers. The sources above and the methodology have the arithmetic.

Read the methodology
  • About 0.24 watt-hours for the median Gemini Apps text prompt in May 2025, by Google’s own measurement, published in August 2025. That counts the chips, the host machine, idle capacity and data centre overhead across Google’s fleet. Counting only the active chips in its most efficient data centres, Google gets 0.10 watt-hours.

  • Google publishes no daily figure and no count of prompts. From its token count, over 3.2 quadrillion tokens in May 2026, our methodology gives about 23 GWh of data centre electricity a day on the Gemini Flash tier, roughly 965 MWh an hour. If every token cost what the Pro reasoning tier charges, it would be about 111 GWh a day, which would take almost all the electricity Google’s data centres used in an average day of 2025.

  • Over 3.2 quadrillion across its products and APIs in May 2026, up from roughly 480 trillion in May 2025 and 9.7 trillion in May 2024, about 330 times in two years. Its model APIs alone ran at approximately 22 billion tokens a minute in July 2026.

  • Per prompt, Google’s figures are low: 0.24 watt-hours and 0.03 g of CO2e for a median text prompt in May 2025, on Google’s market-based accounting, which counts its clean energy purchases. At Google’s scale the total is large. On our methodology, the 3.2 quadrillion tokens of May 2026 come to about 333,000 tonnes of CO2e on the Gemini Flash tier, using Google’s own location-based grid factor.

  • The published figures cannot settle it. OpenAI’s chief executive put an average ChatGPT query at 0.34 watt-hours in June 2025 without describing a method. Google measured 0.24 watt-hours for a median Gemini text prompt in May 2025 and published its method. The two use different definitions and boundaries, and neither company publishes the prompt sizes behind its figure.

  • Yes. Google’s Vertex AI bills cached tokens at 10% of the standard input price on Gemini 2.5 and later models. A cached read skips recomputing the prompt, so our methodology charges it a tenth of the energy of an uncached input token. If 70% of the input in Google’s May 2026 month had been read from cache, the Flash-tier estimate would fall by 14.9%, about 50,000 tonnes.

  • The 0.10 watt-hours counts only the active chips, and only for prompts served in Google’s 10% most efficient data centres. The 0.24 watt-hours adds the host CPU and memory, idle machines and cooling overhead, averaged across the whole fleet. Counting the rest in the same efficient data centres gives 0.17 watt-hours; moving to the fleet average gives 0.24 watt-hours.

See your own number instead of ours.

However your team already runs AI, that is how you connect. Paste a read-only provider key, drop in an SDK, add the MCP server, install the CLI or switch on the browser extension, and the first measured tokens reach your dashboard within seconds.

Create your free account

Free plan · first tokens in seconds