cs.AISep 28, 2026

PowerBench: A Benchmark for Agentic Retrieval and Reasoning in Power Systems

Authors: Xijing Wang, Yinsheng Yao, Jinru Ding, Yidong Jiang, Ziwen Xu, Yiwen Jiang, Jie Xu, Dawei Cheng

Organizations: Tongji University · Johns Hopkins University · Shanghai Artificial Intelligence Laboratory · Carnegie Mellon University · National University of Singapore · Big Data Center of State Grid Corporation of China

Abstract

Large language model (LLM) agents offer new opportunities for automated analysis in industry. However, rigorous evaluation of such agents-for example, within power system scenarios-remains hindered: real operational data are confidential, and existing public resources fail to fully capture the chained dependencies and heterogeneous evidence. To address this gap, we propose PowerBench, comprising (1) a generation framework that derives interconnected heterogeneous operational data through a common dependency chain, and (2) a synthetic dataset generated by this framework. The dataset covers 761 devices across 100 device types, with 13.35 million hourly telemetry records spanning two years and 24,939 operational documents. Building on this dataset, we construct 300 questions across three task families that evaluate frontier LLMs' ability to complete analysis tasks that require autonomous evidence retrieval and reasoning across interconnected and heterogeneous data under restricted tool calls and time budgets. Results demonstrate that the evaluated frontier LLMs remain challenged on these tasks: the best model reaches only 74.2% joint accuracy. Our trace analysis further reveals that model performance varies across evidence discovery, content retrieval, tool use, reasoning over evidence, and answer submission. These findings provide detailed insights for evaluating LLM agents and guiding their reliable deployment in industry. The framework, dataset, and benchmark tasks are available at https://github.com/open-compass/PowerBench.

Figures & tables

Appendix figures & tables4 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. RestoreBench: Can AI Agents Restore Power Flow Convergence?

    Aug 31, 2026Riccardo Mansutti, Andrea Pomarico, Robert Jakob +3Power SystemsOptimal Power Flow

  2. Knowledge Boundary Probing and Demand-Guided Intervention for LLM-Based Power System Code Generation

    May 29, 2026Hui Wu, Xiaoyang Wang, Zhong FanCode Generation