Beyond Energy: When Sustainability Dimensions Reshape LLM Serving Decisions
Organizations: Elmore Family School of Electrical and Computer Engineering, Purdue University, USA
Abstract
Large language model (LLM) serving has environmental impacts across energy consumption, carbon emission, water consumption, and biodiversity loss. Yet these dimensions are largely evaluated in isolation, leaving it unclear when and how they lead to different optimization decisions. We present PRISM, a unified framework for characterizing and optimizing LLM serving across energy, carbon, water, and biodiversity impacts. Our analysis reveals a fundamental distinction: computing configurations determine energy consumption, whereas where and when LLM serving is deployed determine its carbon, water, and biodiversity impacts. Under a fixed deployment choice and operational-only accounting, all dimensions preserve the same energy-based configuration ranking. Deployment rankings can diverge across dimensions, while embodied impacts can break configuration invariance when they exceed a lifecycle crossover boundary. PRISM identifies these conditions, quantifies cross-dimensional regrets, and balances the four dimensions. In regional-routing experiments, PRISM reduces median worst-case regret by 50.2% relative to the strongest baseline.
Figures & tables
| Dimension | Effective Operational Intensity | Embodied Component | Output |
| Energy | — | kWh | |
| Carbon | kg CO 2 e | ||
| Water | m 3 world-eq | ||
| Biodiversity | species year |
Appendix figures & tables36 assets
Supplementary material from the paper’s appendix.
Appendix
| Notation | Meaning | Notation | Meaning |
| Energy, carbon, water, and biodiversity dimensions. | Operational impact intensity of dimension per unit IT energy. | ||
| Sustainability-dimension index, . | Power Usage Effectiveness. | ||
| Workload context and functional unit. | Grid carbon intensity. | ||
| Computing configuration, including the model, GPU type, and parallelism. | Direct datacenter water use per unit IT energy. | ||
| Deployment-region index. | Electricity water-intensity factor. | ||
| Execution-time or routing-interval index. | Water stress factor at location ; is the datacenter-region value. |
| Notation | Meaning | Notation | Meaning |
| Fraction of request group assigned to region . | Request-equivalent count, GPU-seconds per request, and arrival hour of request group . | ||
| Maximum-regret variable in the offline routing Linear Programming. | Rolling-horizon forecast length in hours. | ||
| Fraction of request group assigned to configuration in region . | Set of agentic request groups. | ||
| Mean completion time of successful agentic tasks under configuration . | Weight of agentic request group in the completion-time objective. | ||
| Mean completion time, its feasible optimum, and completion-time regret. | Environmental impact, independent optimum, and regret for dimension in the agentic extension. | ||
| Total agentic-group weight, . | Candidate deployment-region set |
| Dimension | Effective Operational Intensity | Embodied Component | Output |
| Energy | — | kWh | |
| Carbon | kg CO 2 e | ||
| Water | m 3 world-eq | ||
| Biodiversity | species year |
| Dataset | Description | Prompt Length | Response Length | ||||
| P50 | P90 | P95 | P50 | P90 | P95 | ||
| ShareGPT (2022) | Open-ended chatbot conversations sampled from real user-LLM interactions. | 31 | 701 | 1,377 | 243 | 568 | 703 |
| RepoBench ( Liu et al., 2024 ) | Repository-level code completion with relevant in-repository context included in the prompt. | 691 | 5,366 | 6,732 | 3 | 7 | 9 |
| LongBench ( Bai et al., 2023 ) | Long-document summarization of government reports. | 8,432 | 17,316 | 21,185 | 655 | 876 | 934 |
| Workload | TTFT | TPOT | Execution Time |
| ShareGPT | 1000 ms | 150 ms | — |
| RepoBench | 5000 ms | 75 ms | — |
| LongBench | s | 150 ms | — |
| SWE-Bench Verified | — | — | 30 min |
| Family | Evaluated Models |
| Llama 3.1 ( Grattafiori et al., 2024 ) | 8B-Instruct, 70B-Instruct |
| GPT-OSS ( OpenAI, 2025 ) | *20B, *120B |
| Qwen3 ( Qwen Team, 2025 ) | 4B-Instruct / Thinking, 8B, 14B, *30B-A3B-Instruct / Thinking, 32B, *235B-A22B-Instruct / Thinking |
| Gemma 4 ( Google DeepMind, 2026 ) | E4B-it, *26B-A4B-it, 31B-it |
| Testbed | Accelerator | Accelerator Memory | CPU / VM | DRAM / Storage | Network |
| L40 | 4 NVIDIA L40 ( NVIDIA, 2022b ) 300 W TDP | 48 GB GDDR6/GPU 864 GB/s | AMD EPYC 7443 ( AMD, ) 24 cores, 200 W TDP 6 CPU cores/GPU | 128 GB DRAM/GPU 1 TB HDD | Intel X550-AT2 ( Intel, ) 10 GbE, single-port operation 6.1 W typical |
| A100 | 4 NVIDIA A100 SXM ( NVIDIA, 2020 ) 400 W TDP | 40 GB HBM2/GPU 1,555 GB/s | AMD EPYC 7763 ( AMD, ) 64 cores, 280 W TDP 16 CPU cores/GPU | 128 GB DRAM/GPU 1 TB SSD | NVIDIA ConnectX-6 ( NVIDIA, ) 100 Gb/s link 19.58 W typical |
| H100 | 8 NVIDIA H100 SXM ( NVIDIA, 2022a ) up to 700 W TDP | 80 GB HBM3/GPU 3.35 TB/s | Intel Xeon Platinum 8480+ ( Intel, ) 56 cores, 350 W TDP 7 CPU cores/GPU | 128 GB DRAM/GPU 1 TB SSD | 10 NVIDIA ConnectX-7 ( NVIDIA, ) 400 Gb/s 24.9 W typical |
| Agent sandbox | – | – | GCP e2-highmem-16 ( Google Cloud, ) 16 vCPUs Intel Xeon E5-2699 v4 ( Intel, ) 22 cores / 44 threads, 145 W TDP | 128 GB DRAM 100 GB SSD 1 TB HDD | GCP virtual network |
| Model | ShareGPT | RepoBench | LongBench | SWE-Bench | |
| ES | EM | Summarization | Verified | ||
| Llama-3.1-8B-Instruct | 13.8 | 43.5 | 8.2 | 36.6 | 0.20 |
| Llama-3.1-70B-Instruct | 17.8 | 56.0 | 24.8 | 37.7 | 0.20 |
| GPT-OSS-20B | 73.7 | – | – | 30.9 | 2.03 |
| GPT-OSS-120B | 112.8 | – | – | 28.8 | 4.46 |
| Qwen3-4B-Instruct | 50.5 | 44.6 | 9.2 | 30.9 | 3.65 |
| Location | Provider | Region code |
| Vienna, Austria | Azure | austriaeast |
| Brussels, Belgium | GCP | europe-west1 |
| Hamina, Finland | GCP | europe-north1 |
| Paris, France | AWS | eu-west-3 |
| Frankfurt, Germany | GCP | europe-west3 |
| Milan, Italy | AWS | eu-south-1 |
| Figure | Low (%) | High (%) | Figure | Low (%) | High (%) |
| Figure 2 | 1.35 | 3.72 | Figure 14 | 2.18 | 3.90 |
| Figure 4 | 2.23 | 2.64 | Figure 16 | 1.54 | 5.30 |
| Figure 5 | 1.83 | 2.51 | Figure 17 | 1.36 | 6.44 |
| Figure 9 | 1.54 | 3.83 | Figure 18 | 1.43 | 4.12 |
| Figure 10 | 1.48 | 2.09 | Figure 19 | 1.35 | 5.89 |
| Figure 11 | 1.43 | 2.23 | Figure 27 | 3.83 | 6.93 |
| Workload / configuration pair | Dimension | Critical lifetime | ||||
| ( ) | (J/req.) | (yr) | ||||
| RepoBench ( Figure 5 ) Qwen3 4B-I Qwen3 30B-I | Carbon | 9.422 | 5.342 | 2.924 | 1.827 | 10.96 |
| ShareGPT ( Figure 27 ) Qwen3 30B-I Gemma 4 26B | Carbon | 0.341 | 5.201 | 2.924 | 1.778 | 10.67 |
| Figure 29 ① H100 TP1 A100 TP1 | Carbon | 4.482 | 8.566 | 2.924 | 2.929 | 17.58 |
| Water | 4.482 | 0.106 | 7.016 | 0.015 | 0.09 | |
| Biodiversity | 4.482 | 0.288 | 1.047 | 0.275 | 1.65 |
| Region set | WaterWise | PRISM | Median reduction |
| Original 6 | 43.5% [28.8, 111.2] | 37.2% [25.5, 53.5] | 19.1% |
| All 12 | 157.8% [55.1, 212.3] | 75.1% [41.4, 104.0] | 50.2% |
| Concentrated 6 | 101.7% [57.7, 208.5] | 68.1% [48.5, 107.4] | 37.9% |
| Oracle forecast | 20% forecast error | |||
| Horizon | Max. regret | Gap | Max. regret | Gap |
| 1 h | 88.014% | 0.697% | 88.014% | 0.697% |
| 6 h | 87.482% | 0.088% | 87.532% | 0.146% |
| 12 h | 87.415% | 0.012% | 87.517% | 0.129% |
| 24 h | 87.405% | % | 87.525% | 0.138% |
| Offline PRISM | 87.400% | – | – | – |