Organizations: Tongji University · Johns Hopkins University · Shanghai Artificial Intelligence Laboratory · Carnegie Mellon University · National University of Singapore · Big Data Center of State Grid Corporation of China
Large language model (LLM) agents offer new opportunities for automated analysis in industry. However, rigorous evaluation of such agents-for example, within power system scenarios-remains hindered: real operational data are confidential, and existing public resources fail to fully capture the chained dependencies and heterogeneous evidence. To address this gap, we propose PowerBench, comprising (1) a generation framework that derives interconnected heterogeneous operational data through a common dependency chain, and (2) a synthetic dataset generated by this framework. The dataset covers 761 devices across 100 device types, with 13.35 million hourly telemetry records spanning two years and 24,939 operational documents. Building on this dataset, we construct 300 questions across three task families that evaluate frontier LLMs' ability to complete analysis tasks that require autonomous evidence retrieval and reasoning across interconnected and heterogeneous data under restricted tool calls and time budgets. Results demonstrate that the evaluated frontier LLMs remain challenged on these tasks: the best model reaches only 74.2% joint accuracy. Our trace analysis further reveals that model performance varies across evidence discovery, content retrieval, tool use, reasoning over evidence, and answer submission. These findings provide detailed insights for evaluating LLM agents and guiding their reliable deployment in industry. The framework, dataset, and benchmark tasks are available at https://github.com/open-compass/PowerBench.
Figures & tables
Resource
Data Heterogeneity
Data Connectivity
Task Complexity
Ground-Truth Verifiability
SGCC
✗
✗
✗
✗
PowerGridQA
✗
✗
✗
✗
Real-E
✗
✗
✗
✗
MultimodalSyntheticPowerGrid
✓
✓
✗
✗
WorkArena
✓
✗
✓
✓
τ -Bench
✗
✗
✓
✓
Table 1: Comparison with representative datasets and benchmarks. Data Connectivity: whether heterogeneous data sources have logical dependency or ordering relations, not just the same information in different forms. Task Complexity: whether the resource defines agentic tasks requiring retrieval and reasoning. Ground‑Truth Verifiability: whether agent outputs and actions can be programmatically validated against physical, logical, or operational constraints.
Figure 1: Overview of PowerBench. All datasets come from the same structured state. Executable rules are used both to build the data and to verify the ground truth. The questions in the benchmark then ask models to retrieve and combine evidence from the world these rules produce, and to reason over that evidence.
Statistic
Value
Device Type
100
Device Instance
761
Metric
1,275+203=1,478
Rule
2,044
Telemetry
13,350,984
Metric Value
207,878,856
Table 2: Summary statistics of PowerBench.
Figure 2: The three task families in PowerBench: temporal, contextual, and comparative retrieval and reasoning. Each task requires the agent to retrieve heterogeneous evidence from the dataset and return a structured JSON answer and an evidence trace.
Name
Purpose
list_corpus
List corpus files
search_corpus
Search paths, text, and headers
read_file
Read Markdown or JSON files
read_timeseries
Materialize CSV windows
python
Compute results
Table 3: Retrieval tools shared by all models.
Model
TRR
CRR
CoRR
Overall
JA
FA
JA
FA
JA
FA
Macro-JA [95% CI]
FA [95% CI]
GPT-5.6-Sol
85.0
88.1
88.9
93.5
48.7
74.0
74.2 [69.5, 78.7]
82.7 [79.4, 85.8]
GLM-5.3
80.0
91.9
96.7
97.4
13.3
35.1
63.3 [59.2, 67.3]
65.2 [62.2, 68.1]
Doubao-Seed-Evolving
86.7
94.4
91.1
95.6
0.0
21.0
59.3 [55.7, 62.6]
58.1 [56.4, 59.6]
DeepSeek-V4-Pro
50.0
68.3
88.9
90.0
0.7
13.8
46.5 [41.7, 51.1]
47.6 [44.5, 50.4]
Gemini-3.6-Flash
55.0
67.2
82.2
90.7
2.0
10.0
46.4 [41.3, 51.3]
45.6 [42.6, 48.7]
Table 4: Main results (in %). Macro‑JA and overall FA include 95% bootstrap confidence intervals. Best values among the evaluated models are marked with dark pink; second‑best values are marked with light pink.
Model
TRR
CRR
CoRR
RER
AER
VS
TR
TE
RER
AER
VS
TR
TE
RER
AER
VS
TR
TE
GPT-5.6-Sol
92.08
92.79
100.00
90.80
17.43
93.52
79.07
100.00
94.77
22.79
56.87
81.21
100.00
91.84
18.32
GLM-5.3
93.83
89.15
98.33
84.44
9.87
81.67
72.96
98.89
81.28
16.37
60.90
65.47
87.33
95.42
14.96
Doubao-Seed-Evolving
92.86
96.46
100.00
80.62
13.55
80.93
74.44
100.00
78.38
17.79
50.25
47.60
100.00
95.44
25.78
DeepSeek-V4-Pro
96.61
95.72
83.33
90.77
8.75
92.22
85.00
91.11
89.21
12.46
63.48
59.43
56.67
98.49
14.36
Gemini-3.6-Flash
88.73
81.95
85.00
84.80
9.16
79.07
76.48
97.78
82.54
14.90
50.23
43.18
53.33
90.34
13.40
Table 5: Statistics by task family for the radar plots. JA and FA are reported in Table 4 . RER: retrieved evidence recall; AER: available evidence recall; VS: schema validity; TR: tool reliability ( 100−TER ); TE: task-wise relative tool efficiency, where the fewest mean tool calls score 100 within each task family. All values are percentages. Best values per column are shaded dark pink and second-best values light pink.
Appendix figures & tables4 assets
Supplementary material from the paper’s appendix.
Appendix
Field
Role
Rule ID
Identifier within type c (e.g., K1, K12)
Target object
Linked Stage A metric or concern
Object class
Common metric / special metric / special concern
Applicability
Applicable(r,ad) : empty = all instances of c ; else AND of (config field ∈ value set) clauses
Inputs
Inputs(r) : atomic and/or derived metric names
Point condition
Whether a single sample may fire; consecutive-point count
Appendix
Table 6: Schema of a type-level rule r∈Rc .
Limit
Value
Serial tool-call budget
50
Retrieval/reasoning time budget
300 s
read_file max lines / call
240
read_timeseries max rows / call
500
Total bytes read / question
1.5 MB
python timeout / max stdout
8 s / 12,000 chars
Appendix
Table 7: Default resource limits per question.
Model
#dev.
JA
FA
Calls
Lat. (s)
TBR
GPT-5.6-Sol
3
60.0
83.3
26.7
188.2
4.0
4
46.0
72.3
30.4
227.7
10.0
5
40.0
66.6
35.6
285.8
48.0
DeepSeek-V4-Pro
3
2.0
20.9
38.1
495.8
96.0
4
0.0
11.4
39.8
496.2
90.0
5
0.0
9.1
40.4
477.5
86.0
Appendix
Table 8: CoRR stratified by number of target devices ( N=761 ; 50 questions per cell). JA/FA/TBR in percent; tool calls and latency are per-question means. Latency includes the no-tool finalization turn.
Model
N
JA
FA
RER
AER
TBR
GPT-5.6-Sol
200
31.7
63.7
57.4
80.6
41.7
400
33.3
66.7
57.3
84.7
41.7
600
36.7
65.8
57.2
81.2
36.7
761 †
50.0
74.6
57.9
83.6
21.7
Gemini-3.6-Flash
200
0.0
7.5
51.1
42.9
90.0
400
0.0
5.2
49.8
39.6
93.3
Appendix
Table 9: Same 60 CoRR questions at increasing visible corpus size N . JA/FA/RER/AER/TBR in percent. † : reused main-batch endpoint.
Large Language Model (LLM) agents increasingly automate multi-step engineering workflows through tool use, interpretation of intermediate results, and iterative planning. Diagnosing and resolving non-convergent power flow cases is a promising yet largely unexplored application, as it requires engineering judgment, experimentation, and decision-making within constrained action spaces. We introduce a benchmark that evaluates these capabilities across multiple LLMs and three architectures: \emph{chatbot}, \emph{single agent}, and \emph{multi-agent} systems. The evaluation covers two power grids and 46 cases per grid, each requiring one or more corrective actions to restore convergence. The benchmark defines the simulation environment, observation and action spaces, and evaluation metrics, providing a reproducible foundation for developing agentic AI systems for power system planning and operation. The code is available at https://github.com/Mansutti081/RestoreBench
Riccardo Mansutti, Andrea Pomarico, Robert Jakob +3
1ETH Zürich · 2Politecnico di Milano · 3Harvard University
Executable evaluation -- checking the consequences of an agent's actions with a program rather than grading its prose -- has become a prominent way to assess tool-using AI agents in software settings. Electric power engineering has not yet had an analogous benchmark: language-model use is still dominated by retrieval and text question answering, while agents acting on power-system artifacts remain mostly academic prototypes. We introduce the Power Systems Agent Benchmark, an executable benchmark for power-engineering agents. An agent receives a structured task and returns a structured solution; a deterministic evaluator recomputes the engineering quantities, checks operational constraints, and returns a feasibility flag, a normalized score, and explicit violations. The benchmark contains 41 task families across eight areas of power engineering, from power flow and protection to stability, microgrids, reliability, power quality, and forecasting. Each task is grounded in a citable source, standard, or documented engineering formulation. To resist contamination, held-out cases are synthesized on demand by per-family generators from private seeds: the construction is inspectable, but the instances remain private. In a reference evaluation with three command-line agents, the strongest score near the compact tier's ceiling, a smaller open model trails, and public and held-out performance are broadly consistent; a separate public-split grid with OpenCode and Aider probes harness effects. The reference evaluation doubles as quality control: unanimous failures flag candidate task or evaluator defects, and it exposed a latent evaluator bug missed by self-consistency checks. The evaluators are compact deterministic surrogates, but the task contract allows their internals to be upgraded to simulator-backed checks without changing how tasks are posed or solved.
Large language models (LLMs) are increasingly used to automate power-system analysis, but many utilities and energy-research labs require on-premise serving for confidentiality, regulatory, reproducibility, and cost reasons. This makes the reliability of open-weight models a deployment issue. We show that first-pass failures in power-system code generation are dominated not by reasoning alone, but by structured API-knowledge boundary errors: hallucinated function names, misused parameters, and mishandled result tables in versioned simulation libraries. We introduce PowerCodeBench, an execution-validated benchmark generator that pairs natural-language operator queries with pandapower code and numerical ground truth; an L0-L3 documentation-driven probing procedure that measures per-model API knowledge profiles; and a boundary-aware intervention that combines query-side API demand estimation with targeted proactive documentation injection and routed reactive correction. On a 2,000-task frozen release, we evaluate ten open-weight LLMs (1.5B-480B parameters) and four commercial mid-tier APIs. The intervention improves every evaluated open-weight model of at least 7B parameters and every commercial API by 32 to 56 accuracy points. Open-weight models in the 70B-120B range match the commercial mid-tier accuracy range, while Llama-3.1-405B and Qwen3-Coder-480B lead the panel. The targeted prompts preserve the full-context accuracy ceiling while using 41% of the prompt-token cost. The result is an accuracy-side, deployment-time path toward reliable on-premise LLM assistance for grid-analysis workflows without fine-tuning or cloud inference.
Hui Wu, Xiaoyang Wang, Zhong Fan
Department of Engineering, University of Exeter · Department of Computer Science, University of Exeter