As the capabilities and ubiquity of Large Language Models (LLMs) grow, so does their environmental footprint. Despite calls for responsible AI, the machine learning community lacks standardised practices for carbon accounting. Our automated literature review of the 5,285 papers accepted to NeurIPS 2025 reveals that reporting of environmental impact is nearly non-existent. To catalyse a shift toward sustainable AI, we define standardised sustainability metrics for evaluating model training efficiency, accompanied by simple heuristics to estimate the carbon cost of LLM inference. We implement these metrics in carbonbenchmark, a drop-in software solution for tracking and reporting emissions. Finally, to combat the pursuit of marginal accuracy gains at disproportionate environmental costs, we formalise the Smallest Model that Achieves the Job' (SMAJ), a framework which challenges the field to prioritise computational efficiency and environmental accountability alongside traditional State-of-the-Art' (SotA) accuracy.
Figures & tables
Figure 1 : Heatmap showing agreement levels between each of the models on extracted data from 200 randomly selected papers. There is strong agreement ( 90%+ ) between the five strongest models. Therefore in order to reduce computational cost (and therefore environmental impact), the smallest of these five models (Gemini 3.1-Flash-Lite) was chosen to conduct the literature review.
Figure 2 : Learning curve and mathematical fits evaluating the computational efficiency of TinyViT trained on the FashionMNIST dataset. Blue curves show the simple exponential model of training accuracy, green shows the delayed exponential model, and red shows the more accurate power-law model. This demonstrates that both the power law and delayed exponential functions can accurately fit the learning curve, allowing k and β values to be used as measures of training efficiency.
Appendix figures & tables6 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 3 : Learning curve and mathematical fits evaluating the computational efficiency of ResNet18 trained on the FashionMNIST dataset.
Figure 4 : Learning curve and mathematical fits evaluating the computational efficiency of Logistic Regression on the SST-2 text classification dataset.
Figure 5 : Learning curve and mathematical fits evaluating the computational efficiency of DistilBERT fine-tuned on the SST-2 dataset.
Figure 6 : Learning curve and mathematical fits evaluating the computational efficiency of XGBoost trained on the Adult tabular dataset.
Figure 7 : Learning curve and mathematical fits evaluating the computational efficiency of FT-Transformer trained on the Adult tabular dataset.
contrast
odds.ratio
SE
z.ratio
p.value
Haiku 4.5 / Gemini 3.1 Pro Preview
1.083
0.130
0.6623202
0.9396
Claude Opus 4.6 / Gemini 3.1 Pro Preview
0.801
0.099
-1.7892710
0.3223
Claude Sonnet 4.6 / Gemini 3.1 Pro Preview
1.068
0.129
0.5426289
0.967
Gemini 3.1 Flash Lite / Gemini 3.1 Pro Preview
1.154
0.138
1.1973763
0.6921
Gemini 3 Flash Preview / Gemini 3.1 Pro Preview
0.500
0.065
-5.3050248
<0.0001
Gemma 3.12b / Gemini 3.1 Pro Preview
1.624
0.190
4.1447230
2e-04
Appendix
Table 1 : Pairwise comparisons of LLMs to the chosen baseline model (Gemini 3.1 Pro Preview) using Dunnett’s test.
Large Language Model (LLM) usage in recent years has become increasingly widespread in the Artificial Intelligence in Education (AIED) community. While LLMs offer unique avenues for learners and educators, using LLMs comes with computational and environmental costs. These costs are mostly hidden due to a lack of standardised procedures to measure and report these impacts. To address this gap, we first conducted a literature review of all papers published as part of the AIED 2025 conference proceedings, determining if and how computational or environmental costs of LLMs are reported. Most projects use LLMs, but few report computational resources used and almost none discuss environmental impacts of LLMs as an ethical concern. To address this lack of standardised reporting practices, we propose an open-source method for systematically measuring and reporting the computational expense of LLMs and environmental impact of running Machine Learning (ML) AIED systems. We provide software solutions to measure the carbon footprint for both local and cloud based hardware. We also provide an easy-to-use formula to calculate the computational expense of frontier LLMs even when the exact number of parameters is not known. Overall, we hope to motivate colleagues to use our method to strive for more transparent reporting of hidden costs of using LLMs in the AIED community.
Sabrina C. Eimler, Lukas Erle, Daniel Flood +5
Institute of Computer Science and Institute of Positive Computing, Ruhr West University of Applied Sciences, Lützowstraße 5, 46236 Bottrop, Germany · Centre for Computational Science and Mathematical Modelling, Coventry2026 University, Coventry, United Kindgom, CV1 5FB · Carnegie Mellon University +1
The carbon footprint of any deployed Large Language Model (LLM) accumulates during inference, where repeated use of the model substantially exceeds the one-time cost of fine-tuning. Yet most efficiency interventions target either pre-training scale or post-hoc compression. We ask whether folding a calibrated, differentiable energy surrogate into the fine-tuning objective can produce inference behavior that gains task accuracy at zero or near-zero carbon cost, a break-even configuration. We propose a joint loss mechanism with a per-model carbon-emission parameter, a linear surrogate over parameter norm, FLOP proxy, and a memory proxy, fit from on-hardware energy profiling. We fine-tune three architecturally distinct families: Gemma-2 2B, Llama-3.1 8B, and Qwen-2.5 14B, and evaluate inference F1 and CO2 emissions on three MMLU subjects: abstract algebra, philosophy, and formal logic. We discover from several outcomes that the carbon term behaves as either harmful interference or beneficial regularization depending on the task structure. We position calibrated carbon-aware fine-tuning as a lightweight, drop-in regularizer with a non-empty but model and task-dependent break-even region. This is an ongoing work, and we will release our codebase soon.
Recent Machine Learning (ML) approaches have shown increased performance on benchmarks at the cost of escalating compute demands. Hardware, algorithmic and carbon optimizations have been proposed to curb energy use and environmental impacts. We estimate the environmental impacts associated with training models documented in the Epoch AI database over the last decade, with a particular focus on impacts associated with Large Language Models and the hardware used to train them. We find that energy use and environmental impacts associated with training ML models have increased exponentially, even when considering impact reduction strategies such as using less carbon intensive electricity mixes or more efficient hardware. Optimization strategies do not mitigate the impacts induced by model training, suggesting rebound effect. We show that the impacts of hardware must be considered over the entire life cycle rather than the sole use phase in order to avoid impact shifting. Our study demonstrates that increasing efficiency alone does not ensure sustainability. There is an urgent need to systematically integrate environmental impacts in NLP evaluation practices to better inform the community and support the use of impact as a feature in research planning and decision making.