physics.soc-phOct 1, 2026

Strategic Governance of AI Models in Earth Science

Authors: Makoto Kelp, Amirhossein Arzani, Patricia Castellanos, Paul Griffiths, Ivan Higuera-Mendieta, Manuel Perez-Carrasco, Viral Shah, Patrick Obin Sturm, +1 more

Abstract

AI foundation models pretrained on weather and climate data are increasingly fine-tuned to Earth science tasks well beyond weather forecasting. Their development and adoption are outpacing the scientific community's ability to evaluate them. These models are judged almost entirely by benchmark skill metrics, which measure how closely a forecast reproduces a reference product but not whether a model represents the physical processes governing the system it predicts. Forecast skill and physical reliability are therefore distinct properties. The distinction is most consequential under the nonstationary conditions of a changing climate for which these models were never trained. We identify five priorities for the physical evaluation of AI models in Earth science from task-specific emulators to foundation models, spanning training data, fine-tuning, behavioral testing, mechanistic interpretability, and output validation. We recommend three activities for the coming decade: 1) open AI-ready evaluation datasets, 2) a shared reporting standard for physics-based evaluation, and 3) a dedicated research program on the safety of these models.

Explore similar work

CardsList
  1. Earth Science Foundation Models: From Perception to Reasoning and Discovery

    May 9, 2026Xiangyu Zhao, Bo Liu, Yuehan Zhang +9AI for ScienceGeospatial Foundation Models

  2. Extreme Weather Bench: A framework and benchmark for evaluation of high-impact weather

    May 1, 2026Amy McGovern, Taylor Mandelbaum, Daniel Rothenberg +6Forecasting BenchmarksWeather Forecasting

  3. The physics of AI weather models

    May 22, 2026George Craig, Tobias Selz, Matthias Beylich +1AI for ScienceRepresentational Similarity Analysis