Beyond State-of-the-Art: Standardising Environmental Impact Metrics for AI Research
Organizations: School of Computing, Australian National University · Commonwealth Scientific and Industrial Research Organisation
Abstract
As the capabilities and ubiquity of Large Language Models (LLMs) grow, so does their environmental footprint. Despite calls for responsible AI, the machine learning community lacks standardised practices for carbon accounting. Our automated literature review of the 5,285 papers accepted to NeurIPS 2025 reveals that reporting of environmental impact is nearly non-existent. To catalyse a shift toward sustainable AI, we define standardised sustainability metrics for evaluating model training efficiency, accompanied by simple heuristics to estimate the carbon cost of LLM inference. We implement these metrics in carbonbenchmark, a drop-in software solution for tracking and reporting emissions. Finally, to combat the pursuit of marginal accuracy gains at disproportionate environmental costs, we formalise the Smallest Model that Achieves the Job' (SMAJ), a framework which challenges the field to prioritise computational efficiency and environmental accountability alongside traditional State-of-the-Art' (SotA) accuracy.
Figures & tables
Appendix figures & tables6 assets
Supplementary material from the paper’s appendix.
Appendix
| contrast | odds.ratio | SE | z.ratio | p.value |
|---|---|---|---|---|
| Haiku 4.5 / Gemini 3.1 Pro Preview | 1.083 | 0.130 | 0.6623202 | 0.9396 |
| Claude Opus 4.6 / Gemini 3.1 Pro Preview | 0.801 | 0.099 | -1.7892710 | 0.3223 |
| Claude Sonnet 4.6 / Gemini 3.1 Pro Preview | 1.068 | 0.129 | 0.5426289 | 0.967 |
| Gemini 3.1 Flash Lite / Gemini 3.1 Pro Preview | 1.154 | 0.138 | 1.1973763 | 0.6921 |
| Gemini 3 Flash Preview / Gemini 3.1 Pro Preview | 0.500 | 0.065 | -5.3050248 | <0.0001 |
| Gemma 3.12b / Gemini 3.1 Pro Preview | 1.624 | 0.190 | 4.1447230 | 2e-04 |