Methodology v1.2Watershed open frameworkIPCC Tier 1 uncertainty

How a token becomes a number an auditor can check.

CrbonFree meters AI usage per call, then converts it to CO₂e in three published layers, with a stated uncertainty band and a pinned factor version. This page holds the maths, the factors and the sources, and a calculator that runs them.

3

layers, from active silicon to full lifecycle

±28%

uncertainty band on every figure, IPCC Tier 1

8

model tiers with published per-token factors

221

tonnes CO₂e per trillion tokens at the medium tier

What the CrbonFree methodology is

CrbonFree methodology v1.2 is a token-level method for estimating the greenhouse gas emissions of AI usage. It charges each input, output and cached token a published number of joules for its model tier, converts energy to CO₂e through the facility’s PUE and grid intensity, then widens the boundary in two steps to cover data-centre operation and the embodied and training carbon of the hardware and models, following the Watershed open framework. Every figure carries an IPCC Tier 1 uncertainty band.

Active factor version v1.2. Page last checked 8 September 2026.

Aligned with Watershed’s open framework

Earlier versions reported active silicon only. Version v1.2 keeps that four-way token engine as the innermost layer and adopts Watershed’s open framework for corporate AI emissions around it, so the boundary matches the method other reporters use. CrbonFree adopts the published framework independently: the alignment is in boundary and layer structure, and Watershed’s own reference factors give a larger per-token figure.

The framework, at Watershed

Three layers, then the band

Each layer contains the one before it. The third is the figure the dashboard, the API and the receipts report; the first two are always available underneath it.

  1. 1

    Active energy

    Each token is charged joules for the work it causes: prefill for input, decode for output, a much smaller figure for a cached read. Joules become kilowatt hours, then CO₂e through the facility’s PUE and the grid’s intensity.

  2. 2

    Facility operation

    Accelerators do not run alone. Host power adds 18 percent and clusters average 30 percent utilisation, so the active figure is scaled up to what the data centre actually draws. This is the operational, Scope 2 boundary.

  3. 3

    Full lifecycle

    Manufacturing the hardware adds 0.02 kg and training the model 0.008 kg per million tokens, amortised over the fleet’s serving life. Both are carried as their own line, never blended in silently. This is the figure reported everywhere.

  4. 4

    Uncertainty

    A 20% activity uncertainty and a 20% factor uncertainty combine by root sum of squares to ±28.3 percent. The bounds travel with the number through the dashboard, the API and every receipt.

layer1 = J ÷ 3,600,000 × PUE × grid
layer2 = layer1 × 1.18 ÷ 0.3
layer3 = layer2 + (0.02 + 0.008) × tokens ÷ 1e6

J = uncached × prefill + cache creation × prefill + cached × cachedJ + output × decode. Bounds = layer3 × (1 ± 0.2828).

Per-tier factors, version v1.2

Each model is matched to a tier by name pattern. Reasoning tiers carry a higher decode figure because extended thinking multiplies output-side compute. OpenAI and Anthropic tiers report against a 1.2 PUE and a 0.42 kg per kWh grid; Google tiers use Google’s published 1.09 fleet PUE and a 0.345 kg per kWh mix.

TierExample familiesPrefill J/tokDecode J/tokCached J/tokPUEGrid kg/kWhkg per 1M, 75/25
Small or distilledGPT-4o mini, GPT-5 nano, Claude Haiku, small embeddings0.10.20.011.20.420.1 kg
Medium or flagshipGPT-4o, GPT-5 mini, Claude Sonnet, large embeddings0.30.50.031.20.420.2 kg
Large or frontierGPT-4, GPT-5, Claude Opus0.50.90.051.20.420.4 kg
Reasoning, extended thinkingo1, o3, o4, Claude 3.7 Sonnet0.51.80.051.20.420.5 kg
Gemini FlashGemini 2.5, 3 and 3.5 Flash, Flash Lite0.080.50.0081.090.3450.1 kg
Gemini Pro, reasoningGemini 2.5 and 3.1 Pro0.352.50.0351.090.3450.4 kg
Gemma, open weightsGemma 2, 3 and 4, CodeGemma0.040.160.0041.090.3450.1 kg
Google embeddingsGemini Embedding, text-embedding-004 and 0050.00400.00041.090.3450.0 kg

The last column is the full-lifecycle figure for one million tokens at a 75/25 input to output split with no caching, the assumption used for aggregate counts. The training factor of 0.008 kg per million tokens is a conservative default pending confirmation and sits below Watershed’s reference value, so lifecycle figures are a floor.

Run the method yourself

Pick a volume, a tier and an input share. The calculator uses the same functions every figure on this site uses, and shows the three layers separately.

100M

1M100M10B

0.3 J per input token, 0.5 J per output token, 0.03 J cached. PUE 1.2, grid 0.42 kg per kWh.

75% input, 25% output

Result, full lifecycle

22.1 kg

band 15.8 kg to 28.3 kg at ±28.3% · 10 kWh active energy

  • Layer 1, active energy4.9 kg
  • Layer 2, facility operation14.4 kg
  • Layer 3, embodied and training2.8 kg
layer1 = J ÷ 3,600,000 × PUE × grid · layer2 = layer1 × 1.18 ÷ 0.30 · layer3 = layer2 + (0.02 + 0.008) kg × tokens ÷ 1e6
Cached input tokens are treated as none here; the product meters them separately at the cached rate. Aggregate counts without a cache breakdown are treated as uncached, the conservative default.

What an auditor asks, answered in one place

Which boundary is this?

Operational Scope 2 for the energy, embodied and training carried as Scope 3, following the GHG Protocol value chain standard and the Watershed framework.

Where do the factors come from?

Per-tier active energy from Patterson et al., grid intensity from Ember, PUE from provider disclosures, cache costs from provider documentation. Each is cited below.

What is excluded?

Network transfer, end-user devices and a gateway’s own routing infrastructure. Water is reported separately by the product. Each exclusion pushes the figure lower, so it is a floor.

How is a version fixed?

The factor version, v1.2 today, is pinned on each billing period and recorded on each signed receipt, so a filed report never moves when factors change.

References

  1. 01
    Estimating GHG Emissions from AI Use: Framework for Corporate-Level Measurement

    Bistline et al., 2026. The corporate AI emissions framework the v1.2 layered boundary follows: operational Scope 2 plus lifecycle Scope 3.

  2. 02
    An Open Framework for AI Emissions Measurement

    Watershed’s write-up of the same framework. The paper and the post describe one method.

  3. 03
    GHG Protocol, Corporate Value Chain (Scope 3) Standard

    The boundary between operational Scope 2 and embodied and training Scope 3.

  4. 04
    IPCC 2006 Guidelines, Volume 1, Chapter 3: Uncertainties

    Tier 1 root sum of squares propagation for combined uncertainty.

  5. 05
    Patterson et al., 2021: Carbon Emissions and Large Neural Network Training

    Per-tier active energy per token, prefill against decode.

  6. 06
    Ember, global electricity carbon intensity

    Grid intensity behind the energy to CO2e conversion.

  7. 07
    Anthropic, prompt caching documentation

    Why a cached input token costs almost no compute.

Four terms this page leans on

Token
The unit a language model reads and writes, roughly three quarters of an English word. Providers bill by the token, and so does carbon accounting for AI, because the energy a request uses scales with the tokens it processes.
Scope 3
The GHG Protocol category for emissions a company causes but does not own, such as purchased services. AI inference bought from a provider sits here for the buyer, which is why an unmeasured AI footprint is a Scope 3 gap.
PUE
Power usage effectiveness: total data-centre energy divided by the energy the computing equipment itself uses. A PUE of 1.2 means cooling and power distribution add 20% on top of the servers.
Uncertainty band
The range a figure is expected to fall in, given how well its inputs are known. This study combines a 20% activity uncertainty and a 20% factor uncertainty by root sum of squares, giving plus or minus 28.3%.

Questions

Questions about the method.

Short answers. The specification in the dashboard has every factor and pattern.

Open the specification
  • Because the energy a request uses scales with the tokens it reads and writes, not with what the provider charged for them. Two models can cost the same per token and draw very different power. A spend-based estimate cannot tell them apart; a token-based one can, which is what makes the figure defensible.

  • Every figure carries a band of plus or minus 28.3%, the root sum of squares of a 20% activity uncertainty and a 20% factor uncertainty, following the IPCC Tier 1 approach. The band is reported next to the number, not hidden in a footnote, and the API returns the bounds with every summary.

  • Nothing changes retroactively. The factor version is pinned on each billing period and recorded on each receipt when it is signed, so a report you filed under v1.2 stays a v1.2 report. New periods use the new version, and the specification in the dashboard is updated with the same release.

  • No. Measurement and climate action are kept separate on purpose. The measured figure is reported unchanged beside any retirement receipt, so an auditor sees the footprint and the action as two facts rather than one net number.

  • Yes. The factors, the three formulas and the uncertainty rule are on this page, the calculator below runs them, and the API and SDKs return the same bounds for your own usage. The OpenRouter case study publishes its dataset so the arithmetic can be checked from the raw counts.

See your own number instead of ours.

However your team already runs AI, that is how you connect. Paste a read-only provider key, drop in an SDK, add the MCP server, install the CLI or switch on the browser extension, and the first measured tokens reach your dashboard within seconds.

Create your free account

Free plan · first tokens in seconds