The Environmental Cost of AI
Energy, Carbon, Water, and the Need for Transparent Compute Decisions
Artificial intelligence is usually experienced as software, but it is delivered through energy-intensive physical infrastructure: chips, servers, cooling systems, power grids, water systems, transmission lines, and data centers. Its environmental impact is therefore not determined by model quality alone. It depends on where compute runs, when it runs, what hardware serves it, how efficiently the work is executed, how the local grid is powered, what cooling technology is used, and how the environmental accounting boundaries are drawn.
The most defensible current evidence shows that data centers are no longer a marginal electricity load. The International Energy Agency estimates that data centers consumed about 415 terawatt-hours of electricity in 2024, roughly 1.5 percent of global electricity demand, and projects that this could more than double to about 945 terawatt-hours by 2030, with AI as the main driver of incremental growth. The United States alone represents about 45 percent of global data center electricity use.
Carbon impact is highly location dependent. A request served where coal or gas dominate the marginal generation carries a very different emissions profile than one served where hydro, nuclear, wind, or solar are abundant. National averages can conceal large regional differences, which is precisely why the EPA maintains its eGRID database of emissions rates and generation mix across grid subregions.
Water impact is also material, often underestimated, and more local than carbon. Data centers consume water directly through cooling and indirectly through electricity generation. Lawrence Berkeley National Laboratory estimated that U.S. data centers directly consumed about 66 billion liters of water in 2023, while their indirect water footprint from electricity use was nearly 800 billion liters.
Per-prompt estimates can be useful but must be communicated carefully. Google reported that the median text prompt in its Gemini apps used about 0.24 watt-hours of electricity, 0.26 milliliters of water, and emitted about 0.03 grams of CO2e under its own measurement methodology. That is a valuable provider-specific measurement, but experts have cautioned against generalizing any single figure across all models, regions, and workloads, especially when it omits indirect water or location-based emissions.
The central reality is this. Individual AI interactions may appear small, but aggregate AI infrastructure growth is large enough to affect grids, water systems, emissions inventories, capital planning, and local permitting. The responsible conversation should move beyond one prompt equals X and toward transparent, location-aware estimates that separate energy, carbon, water, uncertainty, and accounting method.
AI Has a Physical Footprint, and It Is Not a Single Number
The environmental footprint of AI is not one figure. It is a system-level outcome produced by several interacting variables.
The first is compute demand: model size, context length, output length, modality, routing, batching, and inference efficiency. The second is hardware efficiency: accelerator type, utilization, idle allocation, host-server energy, memory, networking, and storage. The third is data center efficiency: power usage effectiveness, cooling technology, cooling load, and facility overhead. The fourth is grid and geography: the local electricity mix, marginal generation, transmission constraints, water stress, and grid congestion. The fifth is the accounting boundary itself, meaning whether a claim includes only accelerator energy, full server energy, data center overhead, indirect electricity water, market-based renewable matching, location-based emissions, embodied hardware emissions, or model development overhead.
A credible analysis therefore avoids false precision. The better framing is not that AI uses exactly some amount of energy, but that AI workload impact can be estimated within a defined boundary, using disclosed assumptions, region-specific grid and water data, and explicit uncertainty ranges.
This matters because the same task can have very different environmental implications depending on routing, hardware, geography, and time. Intelligent compute placement, smaller-model selection, batching, caching, model distillation, and region-aware routing can reduce estimated energy and emissions without changing the user-facing result.
Why This Issue Has Become Urgent
Data centers are becoming a major category of electricity-demand growth. The IEA expects them to account for roughly one tenth of global electricity demand growth to 2030, and more than 20 percent of demand growth in advanced economies.
The U.S. picture is sharper still. Lawrence Berkeley National Laboratory estimated that U.S. data centers consumed roughly 176 terawatt-hours in 2023, up from 76 terawatt-hours in 2018 and about 60 terawatt-hours in the 2014 to 2016 period. That means U.S. data center electricity use more than doubled between 2018 and 2023. By 2028, its scenarios place demand between about 325 and 580 terawatt-hours, which at the high end would be roughly 12 percent of total U.S. electricity consumption.
AI also changes the infrastructure profile of the internet. AI accelerators concentrate far more power per rack than conventional servers, which raises cooling intensity and often requires advanced air or liquid cooling. Inference demand can be bursty while training runs continuously at high utilization, hardware turns over quickly enough to increase embodied emissions, and compute clusters where power, fiber, land, and cooling are available. The IEA notes that a typical AI-focused data center can use as much electricity as 100,000 households, and the largest now under construction can use 20 times that.
Grid constraints are now a limiting factor in their own right. The question is not only total energy, but whether grids can absorb concentrated new loads without delaying decarbonization, increasing fossil dispatch, or raising local prices. The IEA estimates that about 20 percent of planned data center projects could face delay risk without grid action, and transmission projects often take four to eight years, creating a timing mismatch between fast data center construction and slow grid expansion.
Key Quantitative Anchors
A handful of figures anchor the rest of this discussion. They are best read as well-sourced high-level estimates, not precise measurements.
Global data center electricity use was about 415 terawatt-hours in 2024, roughly 1.5 percent of global electricity, and is projected to reach about 945 terawatt-hours by 2030 in the IEA base case. U.S. data center electricity use was about 176 terawatt-hours in 2023, roughly 4.4 percent of national electricity, with 2028 scenarios spanning about 325 to 580 terawatt-hours, or roughly 6.7 to 12.0 percent of U.S. electricity.
On water, U.S. data centers directly consumed about 66 billion liters in 2023, with an indirect electricity-related footprint of nearly 800 billion liters and a national average indirect water intensity near 4.52 liters per kilowatt-hour. On carbon, U.S. data center electricity-related emissions were about 61 billion kilograms of CO2e under the LBNL grid-mix method, while global data center electricity emissions sit near 180 million tonnes of CO2 today and could approach 300 million tonnes by 2035 in the IEA base case.
For a single interaction, Google reported about 0.24 watt-hours, 0.26 milliliters of water, and about 0.03 grams of CO2e for its median Gemini apps text prompt. That is a useful provider-specific measurement, not a universal AI footprint.
Energy: What AI Consumes and Why
AI energy use is often discussed as a single workload. It is not. Training is the process of creating or updating a model, which can run large clusters of accelerators for days, weeks, or months. It is capital-intensive and highly visible, but episodic. Inference is the process of serving outputs to users. It is recurring and demand-driven, and for widely used models it can dominate total operational impact because every request creates incremental compute. Model development, meaning experiments, failed runs, architecture search, tuning, evaluation, and synthetic data generation, is frequently omitted from public disclosures entirely.
A 2025 study found that while many developers disclose power or carbon for final training runs, far less is known about development, hardware manufacturing, or total water use. For a family of language models between 20 million and 13 billion active parameters, the authors estimated 493 metric tons of CO2e and 2.769 million liters of water, with model development contributing roughly half of the training impact.
Inference energy depends on far more than the model name. It varies with input and output token counts, context-window length, parameter count and architecture, whether routing is dense or mixture-of-experts, batch size, accelerator utilization, cache hits, quantization, speculative decoding, idle allocation, host CPU and memory energy, networking, storage, power usage effectiveness, cooling overhead, and the carbon and water intensity of the local grid. This is why two requests to the same application can have different footprints, and why two models doing the same task can produce materially different outcomes.
For public communication, the most important discipline is to define every variable. Energy per request can be approximated as the sum of accelerator energy, host and server energy, idle allocation, and network and storage allocation, divided across the requests served, then multiplied by the facility power usage effectiveness. Carbon then follows from multiplying energy by a grid emissions factor, and water is the sum of direct cooling water and indirect electricity-generation water. A claim that counts only accelerator energy will always look far lower than one that counts full server energy, facility overhead, and grid-related water. Google’s 2025 inference paper is notable precisely because it argues for full-stack measurement that includes host energy and idle capacity, not active chip power alone.
Carbon: Why Location and Accounting Method Matter
For deployed systems, operational carbon comes largely from electricity consumption during use. The same workload can therefore have very different emissions depending on whether it runs in a coal, gas, hydro, nuclear, wind, or solar heavy region. A kilowatt-hour is not environmentally identical everywhere, which is the entire reason the EPA eGRID database exists.
A major source of confusion is the difference between location-based and market-based emissions. Location-based emissions estimate impact using the actual grid mix where electricity is consumed, which best reflects physical grid impact. Market-based emissions reflect contractual instruments such as renewable energy certificates and power purchase agreements, which is useful for corporate accounting but can obscure the physical footprint of compute in fossil-heavy regions if presented alone. The GHG Protocol Scope 2 Guidance standardizes both, and the responsible approach for AI reporting is to show both where possible, with location-based as the physically meaningful primary number.
Renewable matching is useful but limited. Google reports that since 2017 it has matched 100 percent of its annual global electricity use with renewable purchases while pursuing 24/7 carbon-free energy, reaching a 64 percent global carbon-free average in its 2024 reporting. But annual matching is not the same as hourly, local carbon-free operation. A request served at 8 p.m. in a gas-heavy region is physically different from one served at noon in a solar-rich region, even if the company matches total annual electricity with certificates.
Efficiency gains also do not automatically reduce total emissions. Google reported a data center power usage effectiveness of 1.10 against an industry average near 1.58, and cited best practices that can cut training energy by up to 100 times. Yet its emissions still rose 13 percent year over year, driven by data center growth and supply-chain emissions. This is the classic rebound problem: unit efficiency improves while total consumption grows. The conclusion is not that efficiency is irrelevant, but that it must be paired with transparent workload management and honest aggregate accounting.
Water: The Overlooked Footprint
Data center water use has two components. Direct water use is consumed on site, primarily through evaporative cooling and humidification. Indirect water use is consumed upstream to generate the electricity the facility draws, especially in thermoelectric power plants. A data center with low on-site water can still have a large indirect footprint if its electricity comes from water-intensive generation. Peer-reviewed work in npj Clean Water estimated U.S. data center water consumption at about 1.7 billion liters per day and found that fewer than one third of operators measured water consumption at the time of study.
The common efficiency metric, Water Usage Effectiveness, divides annual site water by IT energy and is typically expressed in liters per kilowatt-hour. It is useful but incomplete, because it excludes the water used to generate electricity. A credible estimate distinguishes on-site cooling water from indirect electricity water, notes whether the water is potable or reclaimed, and flags whether the facility sits in a water-stressed basin.
Scale matters here. LBNL estimated about 66 billion liters of direct U.S. data center water in 2023 and nearly 800 billion liters of indirect water from electricity. In other words, indirect water can be an order of magnitude larger than on-site cooling water, depending on the electricity source and accounting method.
Water impact is also more local than carbon. Carbon dioxide mixes globally, but water scarcity is local, and one liter consumed in a stressed basin is not equivalent to one consumed where water is abundant. Because data centers cluster near power, incentives, and fiber rather than near water, reporting should avoid presenting water as a single global average. A 2023 paper estimated that training GPT-3 in U.S. data centers could directly evaporate about 700,000 liters of freshwater, and projected global AI water withdrawal of 4.2 to 6.6 billion cubic meters by 2027. These are modeled scenarios rather than direct measurements, but they show why water belongs in the discussion alongside energy and carbon.
Why Estimates Differ So Much
Public AI footprint estimates can differ by orders of magnitude because they answer different questions. The largest cause is the accounting boundary. A low estimate might include only accelerator energy, active inference time, a median text prompt, one provider’s optimized stack, annual renewable matching, and on-site cooling water. A higher estimate might add host-server energy, idle capacity, memory and networking, facility overhead, indirect water, location-based emissions, embodied hardware emissions, model development, and transmission losses. Both can be true within their boundary. The problem arises when the boundary is not disclosed.
A second cause is the difference between average and marginal impact. Average grid emissions use the annual generation mix in a region, while marginal emissions estimate which plant responds to incremental demand at a given hour. For routing decisions, marginal factors are often more relevant. A serious methodology discloses which it uses.
A third cause is workload type. A median short text prompt is not comparable to long-context analysis, code generation, agentic multi-step workflows, image or video generation, speech synthesis, batch document processing, fine-tuning, or training. The footprint of AI increasingly depends on the kind of work, not on the word AI as a category.
A fourth cause is provider-specific optimization. A highly tuned inference stack can produce very low per-request energy, but a figure like 0.24 watt-hours reflects one company’s hardware, software, utilization, and accounting choices. Provider measurements are valuable, but they should not be generalized across all models, regions, modalities, or serving stacks.
The Efficiency Metrics, and Their Limits
Power Usage Effectiveness divides total facility energy by IT equipment energy. A value of 1.10 means that for every kilowatt-hour used by IT equipment, another 0.10 is used for cooling, power distribution, and other overhead. It has improved substantially in hyperscale facilities, but it captures none of server utilization, hardware manufacturing emissions, water use, carbon intensity, or whether clean energy is available at the time of consumption. LBNL found U.S. data center infrastructure energy fell from about 40 percent of total use in 2014 to about 30 percent in 2023, even as overall demand climbed.
Water Usage Effectiveness divides annual site water by IT energy. It is important but incomplete, because it usually excludes power-generation water, which is why researchers recommend reporting source-level water alongside facility-level metrics.
Carbon Usage Effectiveness divides total energy-related emissions by IT energy. It is useful only when the emissions factor is clearly defined, because a value based on annual market-based renewable matching is not equivalent to one based on hourly location-based grid emissions.
Regulation and Policy Are Tightening
Europe is moving toward mandatory disclosure. The EU Energy Efficiency Directive requires owners and operators of data centers with at least 500 kilowatts of installed IT power to publish specified information annually, covering energy consumption, power utilization, temperature setpoints, waste heat use, water use, and renewable energy. Facilities above 1 megawatt are expected to reuse waste heat where it is technically and economically feasible, or to show that it is not.
Even where rules exist, disclosure remains incomplete. A 2025 investigation reported that EU reporting has been weakened by confidentiality provisions and that only a minority of eligible data centers submitted the required 2024 information. This reinforces a broader point: AI infrastructure impact is hard to verify independently, because facility-level energy data, model-level energy data, and per-request routing data are often unavailable.
International pressure is increasing as well. In 2026 the UN Secretary-General called for AI companies to disclose carbon, water, and land impacts, proposed an AI environmental transparency initiative, and urged companies to power data centers with renewables by 2030. The direction of travel is clear. AI infrastructure will face more scrutiny, more disclosure expectations, and more pressure to demonstrate credible environmental management.
The Case for Transparent Compute Decisions
AI sustainability discussion tends to focus on building more efficient models. That is necessary but incomplete. Once several models can perform a task adequately, the environmental question becomes operational. Which model, served from which region, on which infrastructure, and at what time, can complete the task with the lowest acceptable energy, carbon, and water impact? This is a new category of optimization, the choice of where and how to compute.
Approached carefully, that choice can reduce impact in several ways: by choosing a smaller model when quality allows, avoiding unnecessary large-model inference, routing to lower-carbon regions when latency and privacy permit, preferring regions with lower water stress, using cached or batched execution, and aligning flexible work with cleaner electricity. The largest near-term savings usually come from not over-using frontier-scale models for tasks that smaller, cheaper, lower-energy models handle well.
Any system that estimates this impact should communicate in ranges and confidence levels rather than false precision. A defensible display might give estimated energy as a band of 0.4 to 1.2 watt-hours, estimated carbon as 0.1 to 0.6 grams of CO2e on a location basis, estimated water including indirect electricity water where available, a likely serving region, a stated confidence level, and a short note on method. That is more credible than claiming a prompt used exactly 0.43 watt-hours and saved exactly 0.19 grams, which looks more polished but is rarely defensible without first-party telemetry and facility-level data.
Uncertainty should be presented as honesty, not weakness. Providers rarely disclose exact serving location per request, workloads are dynamically routed, architectures are often proprietary, utilization varies, grid emissions change hourly, and water intensity depends on cooling system and electricity source. The right posture is to provide transparent estimates using the best available data, disclose the assumptions, and update the method as better telemetry and public reporting become available.
Practical Steps for Companies Using AI
The first principle is to stop treating all AI calls as equal. Simple classification and short-text extraction can run on small models. Routine summarization rarely needs a frontier model. Long technical or legal analysis, high-value code generation, and agentic workflows justify larger models or careful orchestration, while image and video generation deserve their own accounting because they are far more compute-intensive. The central discipline is right-sizing: use the smallest reliable model for the task.
The second principle is to track aggregate usage, not just per-request averages. Organizations should monitor monthly energy, location-based emissions, and water estimates, along with request and token volume by model, the share of requests routed to smaller models, and the share served in lower-carbon regions. The operational risk is that per-request efficiency improves while total usage grows faster.
The third principle is to keep carbon and water separate. The lowest-carbon region is not always the lowest-water region. Hydropower can be low-carbon but hydrologically sensitive, thermoelectric power can be water-intensive, and evaporative cooling can cut electricity use while raising on-site water. Routing should not collapse everything into a single green score unless the weighting is transparent. Better reporting shows energy, carbon, and water separately.
The fourth principle is to use renewable claims carefully. Renewable procurement, annual matching, hourly carbon-free energy, location-based emissions, and market-based emissions are different statements. A strong public position is that renewable procurement matters, but does not replace the need to reduce total electricity demand, route workloads intelligently, and account for local grid and water impacts.
Sources and Further Reading
The figures above draw on the strongest available public sources. The International Energy Agency report Energy and AI, 2025, is the primary reference for global electricity demand, projections, and grid constraints. The Lawrence Berkeley National Laboratory 2024 United States Data Center Energy Usage Report, prepared for the U.S. Department of Energy, is the main source for U.S. electricity, water, and emissions estimates.
For regional carbon, the EPA Emissions and Generation Resource Integrated Database, eGRID, provides grid emissions and generation mix factors, and the GHG Protocol Scope 2 Guidance defines location-based and market-based emissions accounting. Provider-specific measurement is drawn from Google’s 2024 Environmental Report and its 2025 paper on measuring the environmental impact of delivering AI at scale.
On water, the npj Clean Water research on data center water consumption and the paper Making AI Less Thirsty by Li and colleagues inform the direct and indirect water discussion. On policy, the EU Energy Efficiency Directive and recent reporting on AI environmental transparency illustrate the tightening disclosure landscape.
AI’s environmental footprint is not a reason to abandon AI. It is a reason to make AI infrastructure measurable, accountable, and optimized.
The public debate often swings between two oversimplifications. One says AI is just software and its footprint is negligible. The other says AI is inherently unsustainable. The evidence supports a more serious conclusion. The per-request impact can sometimes be small, but aggregate demand is growing fast enough to matter for electricity systems, emissions trajectories, water stress, and infrastructure planning.
The next generation of responsible AI infrastructure should make environmental impact visible at the point of the compute decision, so that energy, carbon, and water become standard optimization dimensions alongside cost, speed, and quality. The question is not whether every prompt is catastrophic or negligible. It is whether companies can make compute decisions transparently enough to reduce unnecessary energy, emissions, and water as AI adoption scales.
Source: IEA Energy and AI, 2025