Every emissions factor for AI gives each token the same energy, but NVIDIA says its newest racks make up to 50 times Hopper’s tokens per watt, so the energy behind a token depends on which chip served it, and no API tells a buyer that. Priced through Crbon Labs methodology v1.2 on three chips and three grids, one billion tokens range from 39 kg of CO2e of electricity on Hopper in Alberta to 0.005 kg on GB300 NVL72 in Québec, and the grid moves the carbon of a token further than the chip does. This post sets NVIDIA’s claims beside independent tests, works the arithmetic for each chip and grid, and explains why one factor per token has to stand for a fleet average until providers say which chip served a request.
Every emissions factor for AI divides energy by tokens. NVIDIA spent 2026 promoting tokens per watt, the output of an AI data centre for each unit of power, which is the same fraction turned over. NVIDIA’s tokens-per-watt claims move that denominator by more than an order of magnitude with each generation. A factor that gives every token the same energy, whatever chip served it and whatever grid powered it, cannot follow those moves. Priced on three generations of NVIDIA hardware and three grids, the same billion tokens spread across nearly four orders of magnitude.
Tokens per watt is energy per token, turned over
NVIDIA’s charts count tokens per second per megawatt. A watt is one joule per second, so tokens per second per watt is tokens per joule, and turned over it is energy per token: joules per token equal one million divided by tokens per second per megawatt. A system that makes a million tokens a second on a megawatt spends one joule on each. Multiply by the grid’s carbon intensity and you have carbon per token. Our post on how much energy AI uses compares the energy per query that providers publish.
Any published tokens-per-watt figure depends on its boundary and its speed. SemiAnalysis, whose benchmarks NVIDIA cites, measures against all-in provisioned power, which includes cooling and the rest of the facility, while other benchmarks count the chips alone. Speed matters as much. A system can serve a few users quickly or many users slowly, and the tokens it makes per megawatt depend on which. At GTC in March 2025 Jensen Huang showed one Hopper megawatt producing 100,000 tokens a second when each user got 100 tokens a second, or 2.5 million when requests were fully batched. That is 10 joules per token or 0.4 joules per token, a 25-times spread on the same hardware from speed alone.
Hopper, GB200 and GB300 are three generations of NVIDIA’s AI hardware
Hopper is the generation of AI chips, called GPUs, that NVIDIA announced in March 2022, and its GPUs are the H100 and the later H200. A GB200 NVL72 is a liquid-cooled rack, one cabinet of servers, that holds 72 GPUs of the Blackwell generation, which NVIDIA announced in March 2024. The 72 GPUs are connected so that the whole rack works as one machine. A GB300 NVL72 is the same kind of rack built with Blackwell Ultra, the upgraded GPU that NVIDIA announced in March 2025. An HGX H100 is a server board that carries up to eight H100 GPUs, and an HGX B200 is one that carries eight of the Blackwell generation’s B200 GPUs.
A Hopper-class token is a token served by Hopper hardware, and a GB200 or GB300 token is one served by a GB200 NVL72 or GB300 NVL72 rack. The text of a token is the same whichever hardware serves it. The electricity spent producing a token differs from one generation of hardware to the next, and NVIDIA’s tokens-per-watt claims compare that electricity per token across Hopper, GB200 and GB300.
NVIDIA’s tokens-per-watt multiples from Hopper to GB300 are best cases, one model at one speed
| Date | Comparison | NVIDIA’s claim | Conditions |
|---|---|---|---|
| 18 March 2024 | GB200 NVL72 against the same number of H100s | Up to 25x lower cost and energy | A 1.8-trillion-parameter model; a projection made before the systems shipped |
| 22 August 2025 | GB300 NVL72 against Hopper | 50x higher AI factory output: 10x faster per user times 5x more throughput per megawatt | Model not stated |
| 9 October 2025 | Blackwell against the previous generation | 10x throughput per megawatt for mixture-of-experts models | SemiAnalysis’s DeepSeek R1 benchmark |
| 16 February 2026 | GB300 NVL72 against Hopper | Up to 50x higher throughput per megawatt | DeepSeek-R1, at the H200’s fastest speed |
| 24 August 2026 | GB300 NVL72 against H200 | Up to 15x, up to 40x and about 80x throughput per megawatt | DeepSeek V4 Pro, DeepSeek-R1 and Kimi K3 |
NVIDIA’s 50x meant revenue opportunity in March 2025 and AI factory output in August 2025, and only from February 2026 throughput per megawatt, the tokens-per-watt meaning. On NVIDIA’s own chart for that last 50x, the gap is about 6 times where the H200 runs most efficiently. The 50x is marked at the H200’s fastest point, where it produces very little. Every multiple in the table is NVIDIA’s best case for one model at one speed.
NVIDIA has also put its claim in carbon terms. In September 2025 it projected 16 kg of CO2e per million tokens for DeepSeek-R1 on HGX H100 and 1.6 kg of CO2e on HGX B200, at 100 tokens a second per user, and it did not publish the emission factor behind either estimate.
Independent tests find 3 to 13 times, not 25 to 50
SemiAnalysis measured an HGX B200 at about 3 times an HGX H100’s tokens per megawatt on gpt-oss 120B in October 2025, and a GB200 NVL72 rack at about 8 times a single H200 node on DeepSeek R1 at 30 tokens a second per user. In September 2026 its agentic benchmark put GB300 NVL72 at 9 to 13 times an H200 at 100 tokens a second per user. ML.ENERGY, which measures GPU energy alone on single machines, found in December 2025 that one B200 used a median 35% less energy per token than one H100 at the same latency.
| Held the same | What changed | Spread in energy or cost per token |
|---|---|---|
| One Hopper megawatt | Speed per user, 100 tokens a second against fully batched | 25x |
| One B200 on DeepSeek V4 Pro | Six weeks of software updates, to June 2026 | 1.7x |
| One GB300 rack on DeepSeek R1 | Multi-token prediction switched on | About 21x in cost |
| One B200 | Batch size | 3x to 5x |
| One GPU type, the B200 | The model | 84x |
Speed, software, batch size and the model each move the energy of a token as far as a generation of hardware does, so a chip’s multiple on its own says little about the carbon in a token.
The same billion tokens ranges 8,000-fold across three chips and three grids
Crbon Labs methodology v1.2 does not name the hardware it assumes. Its medium tier charges 0.3 joules per input token and 0.5 joules per output token of active compute, and at the data centre that lands within about 12% of Epoch AI’s H100-based estimate for a ChatGPT answer, so we read it as Hopper-class. For the newer chips we divide those joules by NVIDIA’s up-to multiples, 25 times for GB200 NVL72 and 50 times for GB300 NVL72, and as a check by the independent multiples, 8 times and 12.6 times. The grids are Alberta at 335 grams of CO2e per kWh and Québec at 2.1 grams per kWh, each a 2024 intensity and preliminary in Canada’s national inventory, and Iowa at 288 grams per kWh, the US Environmental Protection Agency’s state average for 2023. The rest of the methodology stays as published: a data centre overhead of 1.2, meaning the building draws 20% more power than its computers, the host and utilisation factors, and the allowance for hardware and training, all listed on the methodology page.
| One billion tokens, kg of CO2e, electricity (full footprint) | Alberta | Iowa | Québec |
|---|---|---|---|
| Hopper | 39.1 (181.6) | 33.6 (160.0) | 0.245 (29.0) |
| GB200 NVL72, NVIDIA’s up to 25x | 1.56 (34.1) | 1.34 (33.3) | 0.0098 (28.0) |
| GB300 NVL72, NVIDIA’s up to 50x | 0.781 (31.1) | 0.671 (30.6) | 0.00489 (28.0) |
| GB200 NVL72, independent 8x | 4.88 (47.2) | 4.19 (44.5) | 0.0306 (28.1) |
| GB300 NVL72, independent 12.6x | 3.10 (40.2) | 2.66 (38.5) | 0.0194 (28.1) |
- kg of CO2e, electricitydarker is larger, on a log scale
- Widest to narrowest7,982x
Crbon Labs Inc. (2026)
The same billion tokens cost 39 kg of CO2e of electricity on Hopper in Alberta and 0.005 kg of CO2e on GB300 NVL72 in Québec, 7,982 times less: 50 times for the chip multiplied by 160 times for the grid. On the independent ratios the spread is still 2,013 times. The carbon in a token depends on the chip that served it and the grid that powered it, and the grid moves it further than the chip does. Alberta and Iowa, though, are only 16% apart, so between two fossil-heavy grids the chip decides.
One factor per token is a fleet average until providers name the chip
A per-token factor that knows neither the chip nor the grid is an average over every chip and every grid, and a single request can sit nearly four orders of magnitude from that average on electricity alone. Speed, software and the model each move the energy of a token between 2 and 84 times on their own.
Our own methodology uses fleet averages today. It assigns one factor per model tier, published and versioned, because no buyer can yet see which chip served a request, and one grid factor per provider family. Its allowance for hardware and training is fixed as well, at 28 kg of CO2e per billion tokens whatever the chip. On a clean grid or an efficient chip that allowance is most of the total, which is why the totals for one billion tokens sit 6.5 times apart while their electricity sits thousands of times apart.
- Electricity, with host power and idle capacity
- Fixed allowance for hardware and training
Crbon Labs Inc. (2026)
The hardware argues for varying that allowance too. NVIDIA’s product footprints put an HGX B200 baseboard at 2,274 kg of CO2e to manufacture against 1,312 kg of CO2e for an HGX H100, and NVIDIA reports 24% less embodied carbon per unit of computation for the B200. A factor can follow the chip and the grid once providers say which ones served the tokens. Until then it stands for a fleet average, and ours is published with every factor and a version number, which each report records.
NVIDIA’s $40 billion per gigawatt is a price, not an efficiency figure
On NVIDIA’s earnings call of 26 August 2026, its chief financial officer said its revenue opportunity per gigawatt of data centre capacity had grown from roughly $18 billion with Hopper to $25 billion with Blackwell and $40 billion with Vera Rubin. That is NVIDIA’s share of what a customer spends to fill a gigawatt with its systems, and Jensen Huang put the whole data centre at about $60 billion per gigawatt. It is not what the gigawatt earns from tokens, and it is not an efficiency figure.
It does show that power is the fixed quantity and that everything is priced against it. NVIDIA’s content per gigawatt rose about 2.2 times from Hopper to Vera Rubin, while the tokens a megawatt produces rose by one to two orders of magnitude over the same generations. The denominator of every per-token factor is moving far faster than the spending behind it.
No framework asks which chip served a request, because no API says
None of the frameworks we checked asks which chip served a request, because no serverless API tells the buyer. Watershed’s open framework for measuring emissions from AI usage, published by Bistline and colleagues in July and August 2026 and adopted by our methodology, works from token volumes, energy per token, the data centre overhead and the grid. It asks providers to disclose energy per token by model and region, refreshed every two quarters, which would capture a change of chip in the result without naming the chip, and its default for embodied emissions assumes an H100-class card. EcoLogits assumes NVIDIA H100s for every model and provider, and the AI Energy Score benchmarks only on H100s, on purpose, so that models can be compared. Jegham and colleagues infer the hardware from outside in a May 2025 study, and a June 2026 preprint by Llopis reports a range from H100 to B200. Google measures its own fleet, which captures the hardware without naming it, and no framework asks a buyer to name it, because a buyer cannot.
What the chip behind your tokens means for your AI emissions
Priced on three chips and three grids, the same billion tokens cost between 39 kg of CO2e of electricity on Hopper in Alberta and 0.005 kg of CO2e on GB300 NVL72 in Québec, and the grid moves the carbon further than the chip does. Every one of NVIDIA’s tokens-per-watt multiples is a best case for one model at one speed, and the independent measurements sit at 3 to 13 times rather than 25 to 50.
For a company that buys tokens through an API, the chip behind its emissions is out of sight. The serverless APIs return the tokens of every call, by model, and they do not say which chip served the call: OpenAI’s response headers carry the organisation, the processing time, the API version and a request ID; Anthropic’s carry a request ID, the organisation and the workspace; OpenRouter records the provider that served a request; and Amazon Bedrock’s global profiles choose the region for you. Anthropic said in September 2025 that it serves Claude on AWS Trainium, NVIDIA GPUs and Google TPUs, and OpenAI agreed in January 2026 to add 750 megawatts of Cerebras systems to serve its customers, so even the maker of the chip is not a given. Only dedicated deployments, where a buyer rents named GPUs, tell you the hardware. The emissions total you can defend comes from tokens by model, converted by a published factor whose version is recorded, and the two things worth asking a provider for are the region that served you and its energy per token, which is what Watershed’s framework asks providers to disclose.
CrbonFree meters those tokens by model from every provider it connects and converts them with Crbon Labs methodology v1.2, factor by factor and version by version, so when a provider does name the chip or the region, the same record can be recomputed rather than estimated again.
Sources
- 01NVIDIA, “NVIDIA Announces Hopper Architecture, the Next Generation of Accelerated Computing”, 22 March 2022
The Hopper generation and its first GPU, the H100; an HGX H100 server board carries four or eight of them.
- 02NVIDIA, “NVIDIA Supercharges Hopper, the World’s Leading AI Computing Platform”, 13 November 2023
The H200, the second GPU of the Hopper generation.
- 03NVIDIA, “NVIDIA Blackwell Platform Arrives to Power a New Era of Computing”, 18 March 2024
GB200 NVL72 against the same number of H100s: up to 25x lower cost and energy, a projection for a 1.8-trillion-parameter model. The rack is liquid cooled and holds 72 Blackwell GPUs that act as a single GPU; an HGX B200 server board links eight B200 GPUs.
- 04Jensen Huang, GTC 2025 keynote, 18 March 2025 (transcript)
One Hopper megawatt making 100,000 tokens a second at 100 per user, or 2.5 million fully batched.
- 05NVIDIA, “NVIDIA Blackwell Ultra AI Factory Platform Paves Way for Age of AI Reasoning”, 18 March 2025
GB300 NVL72: 72 Blackwell Ultra GPUs in one rack, built on the Blackwell architecture; a 50x revenue opportunity against Hopper.
- 06NVIDIA Technical Blog, “Inside NVIDIA Blackwell Ultra: The Chip Powering the AI Factory Era”, 22 August 2025
GB300 NVL72: 50x AI factory output, 10x per-user speed times 5x throughput per megawatt, against Hopper.
- 07NVIDIA Technical Blog, “NVIDIA HGX B200 Reduces Embodied Carbon Emissions Intensity”, 19 September 2025
DeepSeek-R1 at 100 tokens a second per user: 16 kg CO2e per million tokens on HGX H100, 1.6 on HGX B200 (projected); 24% less embodied carbon per exaflop.
- 08NVIDIA, product carbon footprint summaries for HGX H100 and HGX B200, 2025
Cradle to gate: 2,274 kg CO2e for an HGX B200 baseboard and 1,312 kg for an HGX H100.
- 09NVIDIA, “New SemiAnalysis InferenceX Data Shows NVIDIA Blackwell Ultra Delivers up to 50x Better Performance”, 16 February 2026
GB300 NVL72: up to 50x throughput per megawatt against Hopper on DeepSeek-R1, marked at the H200’s fastest point.
- 10NVIDIA Technical Blog, “NVIDIA Vera Rubin and Blackwell Set a New Standard for Agentic AI Performance per Watt”, 24 August 2026
GB300 NVL72 against H200: up to 15x on DeepSeek V4 Pro, up to 40x on DeepSeek-R1, about 80x on Kimi K3.
- 11NVIDIA, second-quarter fiscal 2027 earnings call, 26 August 2026 (transcript)
Revenue opportunity per gigawatt: about $18 billion with Hopper, $25 billion with Blackwell, $40 billion with Vera Rubin.
- 12SemiAnalysis, “InferenceMAX: Open Source Inference Benchmarking”, 9 October 2025
HGX B200 about 3x HGX H100 on gpt-oss 120B; GB200 NVL72 about 8x a single H200 node on DeepSeek R1 at 30 tokens a second per user; per all-in megawatt.
- 13SemiAnalysis, “Rubin NVL72 Agentic Inference”, 14 September 2026
At 100 tokens a second per user on DeepSeek V4 Pro: GB300 NVL72 at 21.1 to 28.5 million tokens a second per megawatt against 2.26 million for H200.
- 14ML.ENERGY, “Diagnosing Inference Energy Consumption with the ML.ENERGY Leaderboard v3.0”, December 2025
One B200 against one H100 at matched latency: a median 35% less GPU energy; batch size alone moves energy per token 3x to 5x.
- 15Environment and Climate Change Canada, National Inventory Report 1990 to 2024, Annex 7 electricity tables, April 2026
Generation intensity in 2024 (preliminary): Alberta 334.757 g CO2e per kWh, Québec 2.097 g.
- 16US Environmental Protection Agency, eGRID2023
Iowa’s state total output emission rate for 2023: 634.156 lb CO2e per MWh, 288 g per kWh.
- 17Epoch AI, “How much energy does ChatGPT use?”, February 2025
An H100-based estimate of about 0.3 Wh for a 500-token GPT-4o answer, the comparison behind our Hopper-class reading of v1.2.
- 18Bistline et al., “Estimating GHG Emissions from AI Use: Framework for Corporate-Level Measurement”, Watershed, August 2026
The framework methodology v1.2 adopts: token volumes, energy per token, PUE and grid, with suggested provider disclosures by model and region and an H100-class card for embodied emissions.
- 19EcoLogits, “LLM Inference” methodology
Assumes NVIDIA H100 GPUs for every model and provider.
- 20AI Energy Score
Benchmarks models only on NVIDIA H100 GPUs, to hold the hardware fixed.
- 21Jegham et al., “How Hungry is AI? Benchmarking Energy, Water, and Carbon Footprint of LLM Inference”, arXiv, 2025
Infers the hardware behind public APIs from their measured performance.
- 22Llopis, “Accounting for AI Inference in Corporate GHG Inventories: A Four-Tier Methodology for Scope 3 Category 1 Reporting”, arXiv, 9 June 2026
Reports a range: H100 for the central estimate, B200 for the lower bound.
- 23Anthropic, “A postmortem of three recent issues”, 17 September 2025
Claude is served on AWS Trainium, NVIDIA GPUs and Google TPUs.
- 24Cerebras, “OpenAI partners with Cerebras to bring high-speed inference to the mainstream”, 14 January 2026
750 megawatts of Cerebras systems to serve OpenAI’s customers, rolling out from 2026.
Figures attributed to CrbonFree come from methodology v1.2, published in full with every factor and formula. Read the methodology.
About the author
Cory Bergh
Cory leads Crbon Labs, which originates its own climate projects and builds CrbonFree, the platform that measures the footprint of AI usage per token. He was previously VP of Technology and Innovation in the energy industry and holds a BComm, an MBA and the CFA.
Share this post
Send it to whoever asked the question.

