Which boundary is this?
Operational Scope 2 for the energy, embodied and training carried as Scope 3, following the GHG Protocol value chain standard and the Watershed framework.
CrbonFree meters AI usage per call, then converts it to CO₂e in three published layers, with a stated uncertainty band and a pinned factor version. This page holds the maths, the factors and the sources, and a calculator that runs them.
3
layers, from active silicon to full lifecycle
±28%
uncertainty band on every figure, IPCC Tier 1
8
model tiers with published per-token factors
221
tonnes CO₂e per trillion tokens at the medium tier
CrbonFree methodology v1.2 is a token-level method for estimating the greenhouse gas emissions of AI usage. It charges each input, output and cached token a published number of joules for its model tier, converts energy to CO₂e through the facility’s PUE and grid intensity, then widens the boundary in two steps to cover data-centre operation and the embodied and training carbon of the hardware and models, following the Watershed open framework. Every figure carries an IPCC Tier 1 uncertainty band.
Active factor version v1.2. Page last checked 8 September 2026.
Earlier versions reported active silicon only. Version v1.2 keeps that four-way token engine as the innermost layer and adopts Watershed’s open framework for corporate AI emissions around it, so the boundary matches the method other reporters use. CrbonFree adopts the published framework independently: the alignment is in boundary and layer structure, and Watershed’s own reference factors give a larger per-token figure.
The framework, at WatershedEach layer contains the one before it. The third is the figure the dashboard, the API and the receipts report; the first two are always available underneath it.
Each token is charged joules for the work it causes: prefill for input, decode for output, a much smaller figure for a cached read. Joules become kilowatt hours, then CO₂e through the facility’s PUE and the grid’s intensity.
Accelerators do not run alone. Host power adds 18 percent and clusters average 30 percent utilisation, so the active figure is scaled up to what the data centre actually draws. This is the operational, Scope 2 boundary.
Manufacturing the hardware adds 0.02 kg and training the model 0.008 kg per million tokens, amortised over the fleet’s serving life. Both are carried as their own line, never blended in silently. This is the figure reported everywhere.
A 20% activity uncertainty and a 20% factor uncertainty combine by root sum of squares to ±28.3 percent. The bounds travel with the number through the dashboard, the API and every receipt.
J = uncached × prefill + cache creation × prefill + cached × cachedJ + output × decode. Bounds = layer3 × (1 ± 0.2828).
Each model is matched to a tier by name pattern. Reasoning tiers carry a higher decode figure because extended thinking multiplies output-side compute. OpenAI and Anthropic tiers report against a 1.2 PUE and a 0.42 kg per kWh grid; Google tiers use Google’s published 1.09 fleet PUE and a 0.345 kg per kWh mix.
| Tier | Example families | Prefill J/tok | Decode J/tok | Cached J/tok | PUE | Grid kg/kWh | kg per 1M, 75/25 |
|---|---|---|---|---|---|---|---|
| Small or distilled | GPT-4o mini, GPT-5 nano, Claude Haiku, small embeddings | 0.1 | 0.2 | 0.01 | 1.2 | 0.42 | 0.1 kg |
| Medium or flagship | GPT-4o, GPT-5 mini, Claude Sonnet, large embeddings | 0.3 | 0.5 | 0.03 | 1.2 | 0.42 | 0.2 kg |
| Large or frontier | GPT-4, GPT-5, Claude Opus | 0.5 | 0.9 | 0.05 | 1.2 | 0.42 | 0.4 kg |
| Reasoning, extended thinking | o1, o3, o4, Claude 3.7 Sonnet | 0.5 | 1.8 | 0.05 | 1.2 | 0.42 | 0.5 kg |
| Gemini Flash | Gemini 2.5, 3 and 3.5 Flash, Flash Lite | 0.08 | 0.5 | 0.008 | 1.09 | 0.345 | 0.1 kg |
| Gemini Pro, reasoning | Gemini 2.5 and 3.1 Pro | 0.35 | 2.5 | 0.035 | 1.09 | 0.345 | 0.4 kg |
| Gemma, open weights | Gemma 2, 3 and 4, CodeGemma | 0.04 | 0.16 | 0.004 | 1.09 | 0.345 | 0.1 kg |
| Google embeddings | Gemini Embedding, text-embedding-004 and 005 | 0.004 | 0 | 0.0004 | 1.09 | 0.345 | 0.0 kg |
The last column is the full-lifecycle figure for one million tokens at a 75/25 input to output split with no caching, the assumption used for aggregate counts. The training factor of 0.008 kg per million tokens is a conservative default pending confirmation and sits below Watershed’s reference value, so lifecycle figures are a floor.
Pick a volume, a tier and an input share. The calculator uses the same functions every figure on this site uses, and shows the three layers separately.
100M
0.3 J per input token, 0.5 J per output token, 0.03 J cached. PUE 1.2, grid 0.42 kg per kWh.
75% input, 25% output
Result, full lifecycle
22.1 kg
band 15.8 kg to 28.3 kg at ±28.3% · 10 kWh active energy
Operational Scope 2 for the energy, embodied and training carried as Scope 3, following the GHG Protocol value chain standard and the Watershed framework.
Per-tier active energy from Patterson et al., grid intensity from Ember, PUE from provider disclosures, cache costs from provider documentation. Each is cited below.
Network transfer, end-user devices and a gateway’s own routing infrastructure. Water is reported separately by the product. Each exclusion pushes the figure lower, so it is a floor.
The factor version, v1.2 today, is pinned on each billing period and recorded on each signed receipt, so a filed report never moves when factors change.
Bistline et al., 2026. The corporate AI emissions framework the v1.2 layered boundary follows: operational Scope 2 plus lifecycle Scope 3.
Watershed’s write-up of the same framework. The paper and the post describe one method.
The boundary between operational Scope 2 and embodied and training Scope 3.
Tier 1 root sum of squares propagation for combined uncertainty.
Per-tier active energy per token, prefill against decode.
Grid intensity behind the energy to CO2e conversion.
Why a cached input token costs almost no compute.
Questions
Short answers. The specification in the dashboard has every factor and pattern.
Open the specificationBecause the energy a request uses scales with the tokens it reads and writes, not with what the provider charged for them. Two models can cost the same per token and draw very different power. A spend-based estimate cannot tell them apart; a token-based one can, which is what makes the figure defensible.
Every figure carries a band of plus or minus 28.3%, the root sum of squares of a 20% activity uncertainty and a 20% factor uncertainty, following the IPCC Tier 1 approach. The band is reported next to the number, not hidden in a footnote, and the API returns the bounds with every summary.
Nothing changes retroactively. The factor version is pinned on each billing period and recorded on each receipt when it is signed, so a report you filed under v1.2 stays a v1.2 report. New periods use the new version, and the specification in the dashboard is updated with the same release.
No. Measurement and climate action are kept separate on purpose. The measured figure is reported unchanged beside any retirement receipt, so an auditor sees the footprint and the action as two facts rather than one net number.
Yes. The factors, the three formulas and the uncertainty rule are on this page, the calculator below runs them, and the API and SDKs return the same bounds for your own usage. The OpenRouter case study publishes its dataset so the arithmetic can be checked from the raw counts.
However your team already runs AI, that is how you connect. Paste a read-only provider key, drop in an SDK, add the MCP server, install the CLI or switch on the browser extension, and the first measured tokens reach your dashboard within seconds.
Free plan · first tokens in seconds