cs.AIJun 3, 2026

Assessing the Carbon Emissions and Energy Consumption of U.S. Hyperscale Data Centers

Authors: Gianluca GuidiFrancesca DominiciTiziano SquartiniCallaway SprinkleJonathan GilmourKevin ButlerEric BellScott Delaney+1 more

Organizations: Department of Biostatistics,Boston,HarvardMassachusetts,T.H. ChanUSASchool of Public Health · Department of Computer Science, University of Pisa, Pisa, Italy · IMT School of Advanced Studies, Lucca, Italy · Environmental Systems Research Institute, Redlands, California, USA · Baxtel, Denver, Colorado, USA[cs.AI] · Department of Environmental Health, Harvard T.H. Chan School of Public Health, Boston, Massachusetts, USA · Department of Biostatistics, UCLA Fielding School of Public Health, Los Angeles, California, USA

Abstract

The rapid proliferation of hyperscale data centers (HDCs) in the US, mainly driven by the adoption of artificial intelligence, has raised concerns about this industry's environmental footprint. We compiled facility-level information on 403 US hyperscale data centers operating between May 2024 and April 2025 and estimated their electricity consumption, electricity sources, and attributable CO2 emissions. Across different facility-load scenarios, these HDCs consumed approximately 68-99 TWh of electricity and were associated with about 37-54 million metric tons of CO2. Under the central scenario, HDC electricity demand corresponded to approximately 1.8% of total US electricity consumption, with roughly 54% of attributed generation supplied by fossil-fuel sources. The HDC electricity-weighted average carbon intensity was approximately 545 gCO2/kWh, about 48% above the contemporaneous US national grid-average carbon intensity of 370 gCO2/kWh. Our approach provides an attributional tool for assessing the environmental footprint of hyperscale data centers using the most recent EPA eGRID plant-level data.

Explore similar work

Jun 5, 2026cs.CY

The Hidden Water Geography of U.S. Hyperscale Data Centers in the AI Era

Water use by data centers is routinely reported as a single footprint, but water is consumed through two physically distinct pathways: at the site for cooling and in the power system that generates electricity. We mapped both pathways for 472 U.S. hyperscale facilities by linking facility locations to electricity regions, hydrologic basins, and water-stress data. Under baseline assumptions, operational water consumption totals approximately 300 GL yr^-1 (range 205-451 across scenarios), with electricity-related water contributing three-quarters of the total. The two pathways produce different hotspot geographies: direct cooling burdens concentrate in stressed western and south-central basins, whereas electricity-related burdens concentrate in a few eastern grid regions with fossil-heavy supply. Just 3 of 24 hosting balancing authorities account for 59% of electricity-related water. Separating pathways identifies which decisions matter where: cooling design and water sourcing locally, electricity planning and procurement regionally
Gianluca Guidi, Francesca Dominici
May 30, 2026econ.EM

Certificates without Electrons? Theory and Evidence on Impacts from AI-Driven Power Demand

Data centers now account for 4.4% of United States electricity demand, yet the grid-level effectiveness of the renewable energy certificates (RECs) and power purchase agreements (PPAs) hyperscalers use to claim carbon neutrality remains unclear. We develop a game-theoretic model in which a data center operator chooses among RECs, PPAs, and behind-the-meter colocation while generators make entry decisions under endogenous financing costs. The model identifies a timing wedge -- the mismatch between consumption and credited renewable generation -- as a central mechanism through which AI demand degrades reliability, raises prices, and increases emissions even when RECs cover 100% of annual consumption. Colocation with storage addresses this wedge directly and induces the greatest renewable entry by eliminating generator revenue risk. We test these predictions by exploiting the staggered release of large language models as a natural experiment, using difference-in-differences on a novel dataset linking AI activity to local grid outcomes. AI demand significantly increases fossil generation, wholesale prices (up to 25% in treated PJM zones), and outage frequency (0.5--1 additional outages per year) near data centers, with impacts scaling in model size. Data centers with on-site generation exhibit a sign reversal in power-quality effects, consistent with the model's prediction that behind-the-meter capacity absorbs demand spikes. Counterfactual analyses show that edge inference, spatial reallocation, and colocated storage each substantially mitigate grid impacts, while REC-only strategies do not. Together, our results demonstrate that the externalities of AI to the grid are tightly coupled to procurement design and the spatial organization of data center infrastructure.
Dana Golden, Aruna Balasubramanian, Niranjan Balasubramanian
Aug 13, 2026cs.DC

InFactPlanner: Planning Sustainable Geo-Distributed LLM Data Centers

The rapid growth of LLM inference is shifting sustainability concerns from one-time training to continuous serving, where infrastructure decisions shape energy use, carbon emissions, water consumption, and service quality. Yet operators often need to compare deployment alternatives before large-scale infrastructure is built, making direct measurement costly, slow, and sometimes infeasible. We present InFactPlanner, a trace-driven decision-support framework for what-if analysis of sustainable AI data center deployment for LLM inference across single and geo-distributed sites. InFactPlanner combines query traces, hardware-model profiles, candidate site configurations, PUE/WUE parameters, renewable generation models, and time-varying grid carbon intensity to estimate power, energy, carbon emissions, water use, latency, and server utilization. The framework abstracts low-level serving effects into configurable hardware-model profiles, enabling rapid comparison of site selection, capacity placement, hardware, model, renewable integration, and routing choices. We validate the energy accounting pipeline by reproducing reference LLM inference energy estimates with less than 10% deviation, evaluate scalability across multiple data centers and server counts, and demonstrate scenario-driven decision analyses for hardware selection, renewable placement, geographic deployment, and carbon-aware routing. Our results show that sustainability-optimal choices can differ from latency-optimal ones, and that the carbon value of deployment depends strongly on the local grid mix.
Nicoletta Tsiopani, Moysis Symeonides, George Pallis +1