Closing Gaps in Emissions Monitoring with Climate TRACE
Authors: Brittany V. Lancellotti, Jordan M. Malof, Aaron Davitt, Gavin McCormick, Shelby Anderson, Pol Carbó-Mestre, Gary Collins, Verity Crane, +27 more
Organizations: Nicholas Institute for Energy Environment & Sustainability, Duke University, Durham, NC, USA. · Department of Electrical Engineering and Computer Science, University of Missouri, Columbia, MO, USA. · WattTime, Oakland, CA, USA. · Johns Hopkins Applied Physics Laboratory, Laurel, MD, USA. · Marine Science Institute, University of California, Santa Barbara, Santa Barbara, CA, USA. · TransitionZero, London, UK. · Global Fishing Watch, Washington, DC, USA. · Duke University, Durham, NC, USA. · Global Fishing Watch, Bainbridge Island, WA, USA. · CTrees, Pasadena, CA, USA.
Global greenhouse gas emissions estimates are essential for monitoring and mitigation planning. Existing emissions datasets provide critical foundations for understanding emissions patterns across sectors, geographies, and time scales. Through a structured assessment of recent emissions datasets, we identified opportunities to further increase the actionability of emissions data through more comprehensive source-level coverage, finer spatial and temporal resolution, and more frequent updates. Building on existing resources to address these opportunities, we present the Climate TRACE framework and resulting dataset, which is available on an open-access platform (climatetrace.org). The Climate TRACE framework synthesizes existing emissions data, prioritizing accuracy, coverage, and resolution, and fills remaining gaps using sector-specific estimation approaches. The resulting dataset is the first to provide global emissions estimates for individual sources (e.g., individual power plants) for most anthropogenic emitting sectors. The dataset spans January 1, 2021, to the present, with a two-month reporting lag and monthly updates. This dataset and open-access platform provides access to detailed emissions estimates for most subnational governments worldwide. By combining source-level spatial detail, monthly updates, and broad sectoral coverage, the dataset is designed to support analyses relevant to emissions monitoring and mitigation planning.
Scope 3 greenhouse gas (GHG) emissions account for the majority of corporate carbon footprints, yet remain difficult to analyze at scale due to sparse disclosures, heterogeneous report document formats, and limited evidence traceability. Existing approaches typically rely on large language models to extract emissions information from ESG reports, but often lack explicit evidence grounding or depend on costly manual annotation and verification to ensure extraction reliability. To address these challenges, we propose Scope3Trace, an evidence-grounded information extraction framework designed to extract interpretable and traceable Scope 3 emissions information from real-world ESG and sustainability reports. The framework integrates a document information extraction pipeline that performs PDF collection and OCR parsing, LLM-assisted page localization and table reconstruction, and hybrid rule-LLM extraction of organization- and building-level emissions disclosures with evidence-grounded verification. Building upon this framework, we further contribute a dual-level, evidence-grounded, multimodal dataset comprising organization-level Scope 3 disclosures extracted from heterogeneous sustainability reports. Scope3Trace enables reliable extraction and transparent integration of heterogeneous sustainability disclosures, achieving high accuracy in extracting Scope 1-3 totals and category-level disclosures from sustainability reports.
Open datasets and benchmarks for entity-level carbon-emission prediction remain fragmented across access, scale, granularity, and evaluation. We introduce GHGbench, an open dataset and benchmark for company- and building-level greenhouse-gas prediction. The company track contains 32,000+ company-year records from 12,000+ firms with Scope 1+2 and Scope 3 disclosures and financial/sectoral signals; the building track harmonises 491,591 building-year records from 13 open sources into a single schema across 26 metropolitan areas (10 U.S., 15 Australian, 1 Singaporean), with climate covariates and multimodal remote-sensing embeddings. GHGbench defines canonical splits with in-distribution and cross-region/city transfer as primary tasks and temporal hold-out plus short-horizon forecasting as supplementary appendix evidence; headline baselines span gradient-boosted trees, a tabular foundation model, MLP, FT-Transformer, and multimodal fusion, with an LLM panel as auxiliary, all evaluated under multi-seed paired-bootstrap tests. Three benchmark-level findings emerge: (i) building emissions are structurally harder than company emissions; (ii) the in-distribution to out-of-distribution gap dwarfs any within-model gap across both the company track and the building track, and a tabular foundation model is, to our knowledge, the first baseline to open a paired-bootstrap-significant gap over tuned trees on a multi-city building-emissions task; (iii) multimodal remote-sensing embeddings help precisely where tabular generalisation breaks. GHGbench also exposes catastrophic city transfer and the sector-factor lookup ceiling as systematic failure modes. Code and reconstruction recipes are available at GHGbench.
Methane is a potent greenhouse gas that significantly contributes to global warming. However, accurately estimating global methane emissions and consumption remains challenging due to the complex interactions among environmental drivers that may vary across spatial and temporal scales. Prior data-driven methods often overlook the inherent spatiotemporal heterogeneity of ecosystems, failing to explicitly capture site-specific characteristics and cross-year evolutionary dynamics. To address these issues, we propose the Contrastive Hierarchical Adaptive Meta-network (CHAM-net), a novel framework that explicitly learns from historical context to capture site-specific dynamics. CHAM-net employs a hierarchical encoder-decoder architecture, in which the encoder captures site-specific characteristics from historical data and then dynamically conditions the decoder to generate the final prediction. Experimental results demonstrate that CHAM-net consistently outperforms all baseline methods on both simulation and observational datasets for methane emission and consumption, achieving nRMSE values as low as 0.43 and 0.88 with corresponding R2 scores up to 0.97 and 0.68 for emission prediction.