Emissions Measurement Is Moving From Estimation to Governance

Companies are adopting artificial intelligence faster than they can measure its environmental effects. Most organizations know what they spend on AI services, but many lack reliable information about token volumes, model-specific electricity use, serving locations, data-center overhead, and training emissions.

A new framework released by Watershed offers a practical way forward. Instead of requiring perfect data before measurement begins, it allows companies to start with the information they possess and improve their estimates as better operational and provider data become available.

The framework’s most important contribution may not be a single emissions number. It is the creation of a management pathway linking emissions measurement to procurement, engineering, cloud architecture, data governance, and sustainability strategy.

Three Tiers Reflect Three Levels of Information

The framework proposes three measurement tiers:

  • Spend Tier: Estimates emissions from vendor spending when operational data are unavailable.
  • Activity Tier: Combines token volumes with modeled electricity and emissions intensities.
  • Provider Tier: Uses provider-reported per-token electricity or emissions information, potentially supplemented by model, region, and infrastructure data.

These tiers can be applied separately to each provider. A company might have Provider Tier information for one vendor, Activity Tier data for another, and only spending information for the remainder. The resulting inventory can therefore reflect differing levels of data quality rather than forcing every vendor into one uniform methodology.

Provider Tier data should not automatically be treated as direct measurement of a customer’s specific workload. Provider disclosures may still represent fleet-wide averages, typical prompts, modeled estimates, or aggregate infrastructure performance. Their value depends on transparency, representativeness, system boundaries, and verification.

Start With Available Data, but Disclose Its Limitations

The framework recommends beginning with the Spend Tier rather than waiting for token-level or provider-specific information. Companies can inventory their AI use, estimate emissions using the best currently available data, request better information, and upgrade their calculations over time.

A spend-based estimate can be documented and reproducible without being operationally precise. Auditability, reproducibility, and accuracy are related but distinct qualities.

Companies should therefore disclose:

  • which tier was used for each provider;
  • which vendors, models, and use cases were included;
  • material assumptions and exclusions;
  • whether training and infrastructure emissions were included;
  • whether emissions factors are location-based or market-based; and
  • how the organization plans to improve the estimate.

More granular data will not necessarily reduce the reported footprint. Depending on provider pricing, utilization, infrastructure efficiency, electricity source, and accounting boundaries, better information could increase or decrease the estimate.

The objective should not be to produce the smallest defensible number. It should be to produce the most decision-useful estimate supported by the available evidence.

Production Systems Can Outperform Isolated Benchmarks

Certain non-production inference benchmarks may substantially overstate energy use in large-scale production systems. Small-scale tests may use batch sizes of one and omit efficiencies such as caching, inference batching, speculative decoding, higher accelerator utilization, and hardware-software co-design.

The Watershed webinar explained that large providers can spread fixed and idle system energy across billions of queries and many machines, producing substantially lower energy use per query than isolated testing may suggest.

This finding should not be generalized to every laboratory study, benchmark, or external estimate. Some analyses use broader boundaries than others. Some exclude data-center overhead, embodied hardware, or training. A production estimate can be more representative of actual deployment while still omitting lifecycle components.

Nor does greater efficiency per query establish that AI’s total electricity demand is declining. Aggregate impact depends on both the energy intensity of each task and the total volume of AI use.

As AI becomes cheaper and more efficient, demand may grow. Efficiency gains can therefore coexist with increasing absolute electricity consumption and emissions.

Google’s Efficiency Gains Show Why System Boundaries Matter

During the webinar, Google described an approximately 33-fold reduction in Gemini energy per prompt attributable to compounded software and hardware improvements. About 23-fold was associated with software changes, including speculative decoding, batching, and mixture-of-experts architectures. Approximately 1.4-fold came from hardware design and higher utilization.

Higher utilization should be understood as using accelerator capacity more intensively, not simply raising physical operating temperatures. Greater utilization can reduce energy per query by spreading fixed and idle system energy across more useful computation.

Watershed’s written materials separately cite a broader Google finding involving a 47-fold reduction in the energy cost of a Gemini prompt. The available information does not establish that the 33-fold and 47-fold figures use identical baselines, denominators, or system boundaries. They should not be treated as interchangeable without additional documentation.

The underlying management lesson is more important than either headline number. AI efficiency is not controlled by one variable. Software architecture, model design, hardware utilization, cooling, networking, serving location, and electricity sourcing interact across the system.

In Google’s reported implementation, these improvements compounded across the stack. That does not mean every intervention will multiply independently in every system. Some measures may overlap, depend on one another, or already be incorporated into the baseline.

Training Emissions Remain Highly Uncertain

Training a frontier model can create a large upfront emissions burden. Allocating that burden to individual tokens or queries requires an estimate of how much the model will be used over its lifetime.

AI providers generally do not disclose lifetime token volumes. That denominator may reveal commercially sensitive information about adoption and business performance.

In the webinar’s illustrative scenario, training represented approximately 30%–60% of the calculated footprint at an assumed 100 trillion lifetime tokens. At an assumed 10 trillion tokens, the training share exceeded 80%. At much larger lifetime usage, the contribution allocated to each query would fall substantially.

These percentages are sensitivity-analysis outputs, not universal measurements of AI emissions.

The framework recommends reporting training emissions separately rather than embedding them in a single operational per-token figure. This preserves visibility into the uncertainty and avoids creating a misleading appearance of precision.

For a customer using an already-trained model, historical training emissions are not an immediate operational lever. They remain relevant to lifecycle accounting, provider selection, model choice, and incentives for future model development.

Serving Region Can Be a Material Lever

Electricity emissions intensity varies significantly by location. During the webinar, speakers noted that U.S. subregional emissions intensity can differ by roughly fivefold.

Where providers permit region selection, companies may be able to reduce operational emissions without retraining the model or redesigning the product.

The opportunity is not universal. Providers may dynamically route queries, customers may lack visibility into actual serving locations, and latency, privacy, security, or data-residency requirements may constrain the available options.

Companies should distinguish among:

  • physical electricity consumption;
  • location-based emissions;
  • market-based emissions;
  • annual renewable-energy matching; and
  • more granular geographic or hourly matching.

A claim that a service is “renewable-powered” may describe contractual energy procurement rather than the physical electricity serving a particular workload at a particular time.

Intensity and Absolute Emissions Answer Different Questions

Block described a common corporate challenge: carbon efficiency can improve while gross emissions continue to rise as the business expands. AI may accelerate that growth without changing the company’s climate commitments.

This distinction should be explicit in AI reporting.

An intensity metric—such as energy or emissions per token—helps evaluate operational efficiency. Absolute emissions show the total environmental burden associated with all activity.

A company can improve one while worsening the other.

Management should therefore track both:

  • total AI-related electricity consumption and emissions; and
  • normalized indicators such as emissions per token, task, transaction, or business outcome.

Token-based metrics support consistent accounting but do not establish functional equivalence among models or tasks. Tokenizers differ, output tokens often require more compute than cached input tokens, reasoning systems may generate internal tokens, and agentic workflows may initiate multiple calls.

A lower per-token footprint does not necessarily mean a model completes the same task more efficiently or produces an equivalent-quality result.

What Companies Should Do Now

1. Inventory AI Use

Identify material vendors, models, applications, business functions, and internally developed systems. Include embedded AI features that may not be managed through a centralized AI procurement process.

2. Map the Available Data

Collect vendor spend, token volumes, model identifiers, serving-region information, cloud data, and provider disclosures. Engineering teams may already track token use for cost and performance management.

3. Assign a Tier by Provider

Do not wait until every vendor can support the same calculation. Apply the best available tier separately and record the evidence supporting each classification.

4. Calculate and Document an Initial Estimate

State the system boundary, emissions factors, allocation rules, exclusions, uncertainty, and reporting period. Separate training estimates from operational emissions.

5. Request Better Provider Information

Ask for model identity, token-use data, electricity intensity, serving location, operational emissions, training methodology, data-center overhead, and material calculation assumptions.

The webinar identified model family, blended per-token electricity intensity, operational emissions, and serving region as essential provider disclosure fields. Additional information could include input-output token splits, training allocations, behind-the-meter electricity, and data-center overhead.

6. Connect Measurement to Decisions

Use the results to inform:

  • model selection;
  • prompt and workflow design;
  • cloud and serving-region choices;
  • vendor procurement;
  • engineering priorities;
  • clean-energy strategy; and
  • sustainability targets.

Measurement creates leverage only when it reaches the teams making technical and commercial decisions. Block emphasized that quantified information gives sustainability teams a more credible basis for engaging product and engineering functions.

From Estimation to Governance

The Watershed framework is a proposed approach, not a settled accounting standard. Important gaps remain around provider transparency, training allocation, embodied hardware, customer-specific infrastructure data, electricity accounting, and comparability among models and workloads.

Those gaps are not a reason to delay measurement.

Companies can begin with documented estimates, state uncertainty openly, request better information, and improve the calculation over time. The purpose is not merely to add another number to the greenhouse-gas inventory.

The larger opportunity is to build an AI governance process that helps companies understand where emissions arise, which decisions can change them, and whether efficiency improvements are keeping pace with the rapid growth of AI use.


Sources: