The environmental impact of deep learning has attracted increasing attention over the past decade. Existing studies mainly focus on the energy and carbon emissions of model training and inference, while the whole development phase is often overlooked. Yet, architecture prototyping and intensive experiments are conducted during this stage, which is highly energy-demanding. In this article, we propose a methodology to estimate these costs, based on activity logs from the Grid5000 shared computing platform used by the LORIA laboratory. As a case-study, we focus on audio projects developed in the Multispeech research team. We evaluate the overall energy cost of four projects, and we compare them to those of training the reported models. Our results show that the energy required for the development phase is 3 to 256 times greater than that required to train the best-performing model alone. These results advocate for a more systematic reporting of energy consumption across the entire life cycle of deep learning-based audio projects.
Figures & tables
Figure 1: Evolution of computational demand (top) and energy consumption (bottom) from 2014 to 2025 for Grid5000 Nancy clusters, with a distinction between CPU-only and GPU servers.
Figure 2: Daily electricity consumption on GPU servers of Multispeech research group on Grid5000 during the 2023–2024 academic year. The red line indicates the mean daily consumption of 260 kWh.
Number
Energy (kWh)
Energy Ratio
Carbon emissions (tCO 2 e)
Jobs
Runs
Best
All
Dev
Dev/All
Dev/Best
France
US
Air travel
Cui et al. 2023 [ 6 ]
1,462
16
11,654
17,042
35,288
2
3
2.8
22.1
4.0
Ayilo et al. 2024 [ 4 ]
337
6
222
624
4,095
7
18
0.3
2.6
4.8
Ayilo et al. 2025 [ 3 ]
3,529
3
37
123
9,462
77
256
0.5
5.9
2.8
Magron et al. 2027 [ 16 ]
1,463
56
456
3,027
24,239
8
53
1.0
12.1
2.2
Table 1: Computational activity, energy consumption, and carbon emissions of the four studied projects. “Jobs” and “Runs” are the total number of jobs, and reported training runs, respectively. “Best”, “All”, and “Dev” refer to the estimated energy for training the best performing model, all reported training runs, and the whole project development, respectively. Carbon emissions are estimated under French and US electricity mixes, and air travel corresponds to a round-trip flight between Paris and the conference location.
Deep-learning speaker verification (SV) increasingly relies on deep neural network backbones, whose environmental impact remains largely undocumented. In this paper, we conduct an evaluation of ResNet architectures trained on VoxCeleb2, varying depth, channel width, and stage distribution, and measure energy consumption and carbon footprint using node-level sensors. Results show a clear point of diminishing returns: deeper or wider models bring only marginal accuracy gains while energy consumption grows steeply. In contrast, mid-sized networks such as ResNet-50 and stage-concentrated variants achieve favorable trade-offs between performance and environmental impact. These findings provide actionable guidelines for designing energy-efficient SV systems.
Hugo Leguillier, Driss Matrouf, Guillaume Lechien +1
Avignon University, LIA, UPR 4128, France · Aday, France
The rise in deployment of large language models has driven a surge in GPU demand and datacenter scaling, raising concerns about electricity use, grid stress, and the impacts of modern AI workloads. Distillation is often promoted as one of the most effective paths to obtain cheaper, more efficient models, yet these claims rarely account for the full end-to-end energy and resource costs, including crucial teacher-side workloads such as data generation, logit caching, and evaluation. We present a comprehensive energy accounting framework that measures the complete computational cost of distillation pipelines via detailed stage-wise tracking of GPU device power consumption. In our experiments, we separate and log empirical energy use across distinct phases and systematically measure the energy and emissions of two common distillation methods: the classic logit-based knowledge distillation and synthetic-data supervised fine-tuning, constructing energy-quality Pareto frontiers that expose the previously ignored costs. From these measurements and analyses, we derive practical design rules for selecting distillation methods and hyperparameters under energy and budget constraints, and release an open-source measurement harness and accounting protocol to provide a standardized foundation for comparable, reproducible distillation research, explicitly accountable for complete pipeline energy impact.
General audio comprehension now covers speech, sound, and music over durations from seconds to hours, driven by large audio-language models (LALMs) that are increasingly omni-modal. Yet the benchmarks that test them still rely on clips of seconds, where scores saturate and models converge; recent long-form efforts extend duration but evaluate long audio much as short clips are. We introduce LongAudioSpan, a benchmark that spans both duration and depth: it pairs audio from 10 minutes to over 2 hours with 3,240 questions across three cognitive levels, namely perception, understanding, and reasoning. Two paths supply the questions, differing in how question content is sourced and how ground truth is obtained. Native QA extracts questions from the audio's content, posing each as a multiple-choice item and an open-ended one graded by detailed rubrics. Anchor QA instead injects ground truth, planting acoustic anchors into the audio and building a perception-to-reasoning chain scored only to the first error. A fully automated pipeline constructs every item through structured captioning, QA generation, and adversarial critic feedback. Evaluating 12 LALMs on LongAudioSpan, we find the hard part comes before reasoning: distilling a few relevant facts from a long, redundant signal. This difficulty grows with audio length and falls hardest on perception, especially temporal grounding. LongAudioSpan is available at https://huggingface.co/datasets/holvan/LongAudioSpan.
Wen Huang, Yunfei Chu, Meng Gao +2
Qwen Team, Alibaba Group · Tsinghua University · The Chinese University of Hong Kong