Organizations: Tongji University · Johns Hopkins University · Shanghai Artificial Intelligence Laboratory · Carnegie Mellon University · National University of Singapore · Big Data Center of State Grid Corporation of China
Large language model (LLM) agents offer new opportunities for automated analysis in industry. However, rigorous evaluation of such agents-for example, within power system scenarios-remains hindered: real operational data are confidential, and existing public resources fail to fully capture the chained dependencies and heterogeneous evidence. To address this gap, we propose PowerBench, comprising (1) a generation framework that derives interconnected heterogeneous operational data through a common dependency chain, and (2) a synthetic dataset generated by this framework. The dataset covers 761 devices across 100 device types, with 13.35 million hourly telemetry records spanning two years and 24,939 operational documents. Building on this dataset, we construct 300 questions across three task families that evaluate frontier LLMs' ability to complete analysis tasks that require autonomous evidence retrieval and reasoning across interconnected and heterogeneous data under restricted tool calls and time budgets. Results demonstrate that the evaluated frontier LLMs remain challenged on these tasks: the best model reaches only 74.2% joint accuracy. Our trace analysis further reveals that model performance varies across evidence discovery, content retrieval, tool use, reasoning over evidence, and answer submission. These findings provide detailed insights for evaluating LLM agents and guiding their reliable deployment in industry. The framework, dataset, and benchmark tasks are available at https://github.com/open-compass/PowerBench.
Figures & tables
Resource
Data Heterogeneity
Data Connectivity
Task Complexity
Ground-Truth Verifiability
SGCC
✗
✗
✗
✗
PowerGridQA
✗
✗
✗
✗
Real-E
✗
✗
✗
✗
MultimodalSyntheticPowerGrid
✓
✓
✗
✗
WorkArena
✓
✗
✓
✓
τ -Bench
✗
✗
✓
✓
Table 1: Comparison with representative datasets and benchmarks. Data Connectivity: whether heterogeneous data sources have logical dependency or ordering relations, not just the same information in different forms. Task Complexity: whether the resource defines agentic tasks requiring retrieval and reasoning. Ground‑Truth Verifiability: whether agent outputs and actions can be programmatically validated against physical, logical, or operational constraints.
Figure 1: Overview of PowerBench. All datasets come from the same structured state. Executable rules are used both to build the data and to verify the ground truth. The questions in the benchmark then ask models to retrieve and combine evidence from the world these rules produce, and to reason over that evidence.
Statistic
Value
Device Type
100
Device Instance
761
Metric
1,275+203=1,478
Rule
2,044
Telemetry
13,350,984
Metric Value
207,878,856
Table 2: Summary statistics of PowerBench.
Figure 2: The three task families in PowerBench: temporal, contextual, and comparative retrieval and reasoning. Each task requires the agent to retrieve heterogeneous evidence from the dataset and return a structured JSON answer and an evidence trace.
Name
Purpose
list_corpus
List corpus files
search_corpus
Search paths, text, and headers
read_file
Read Markdown or JSON files
read_timeseries
Materialize CSV windows
python
Compute results
Table 3: Retrieval tools shared by all models.
Model
TRR
CRR
CoRR
Overall
JA
FA
JA
FA
JA
FA
Macro-JA [95% CI]
FA [95% CI]
GPT-5.6-Sol
85.0
88.1
88.9
93.5
48.7
74.0
74.2 [69.5, 78.7]
82.7 [79.4, 85.8]
GLM-5.3
80.0
91.9
96.7
97.4
13.3
35.1
63.3 [59.2, 67.3]
65.2 [62.2, 68.1]
Doubao-Seed-Evolving
86.7
94.4
91.1
95.6
0.0
21.0
59.3 [55.7, 62.6]
58.1 [56.4, 59.6]
DeepSeek-V4-Pro
50.0
68.3
88.9
90.0
0.7
13.8
46.5 [41.7, 51.1]
47.6 [44.5, 50.4]
Gemini-3.6-Flash
55.0
67.2
82.2
90.7
2.0
10.0
46.4 [41.3, 51.3]
45.6 [42.6, 48.7]
Table 4: Main results (in %). Macro‑JA and overall FA include 95% bootstrap confidence intervals. Best values among the evaluated models are marked with dark pink; second‑best values are marked with light pink.
Model
TRR
CRR
CoRR
RER
AER
VS
TR
TE
RER
AER
VS
TR
TE
RER
AER
VS
TR
TE
GPT-5.6-Sol
92.08
92.79
100.00
90.80
17.43
93.52
79.07
100.00
94.77
22.79
56.87
81.21
100.00
91.84
18.32
GLM-5.3
93.83
89.15
98.33
84.44
9.87
81.67
72.96
98.89
81.28
16.37
60.90
65.47
87.33
95.42
14.96
Doubao-Seed-Evolving
92.86
96.46
100.00
80.62
13.55
80.93
74.44
100.00
78.38
17.79
50.25
47.60
100.00
95.44
25.78
DeepSeek-V4-Pro
96.61
95.72
83.33
90.77
8.75
92.22
85.00
91.11
89.21
12.46
63.48
59.43
56.67
98.49
14.36
Gemini-3.6-Flash
88.73
81.95
85.00
84.80
9.16
79.07
76.48
97.78
82.54
14.90
50.23
43.18
53.33
90.34
13.40
Table 5: Statistics by task family for the radar plots. JA and FA are reported in Table 4 . RER: retrieved evidence recall; AER: available evidence recall; VS: schema validity; TR: tool reliability ( 100−TER ); TE: task-wise relative tool efficiency, where the fewest mean tool calls score 100 within each task family. All values are percentages. Best values per column are shaded dark pink and second-best values light pink.
Appendix figures & tables4 assets
Supplementary material from the paper’s appendix.
Appendix
Field
Role
Rule ID
Identifier within type c (e.g., K1, K12)
Target object
Linked Stage A metric or concern
Object class
Common metric / special metric / special concern
Applicability
Applicable(r,ad) : empty = all instances of c ; else AND of (config field ∈ value set) clauses
Inputs
Inputs(r) : atomic and/or derived metric names
Point condition
Whether a single sample may fire; consecutive-point count
Appendix
Table 6: Schema of a type-level rule r∈Rc .
Limit
Value
Serial tool-call budget
50
Retrieval/reasoning time budget
300 s
read_file max lines / call
240
read_timeseries max rows / call
500
Total bytes read / question
1.5 MB
python timeout / max stdout
8 s / 12,000 chars
Appendix
Table 7: Default resource limits per question.
Model
#dev.
JA
FA
Calls
Lat. (s)
TBR
GPT-5.6-Sol
3
60.0
83.3
26.7
188.2
4.0
4
46.0
72.3
30.4
227.7
10.0
5
40.0
66.6
35.6
285.8
48.0
DeepSeek-V4-Pro
3
2.0
20.9
38.1
495.8
96.0
4
0.0
11.4
39.8
496.2
90.0
5
0.0
9.1
40.4
477.5
86.0
Appendix
Table 8: CoRR stratified by number of target devices ( N=761 ; 50 questions per cell). JA/FA/TBR in percent; tool calls and latency are per-question means. Latency includes the no-tool finalization turn.
Model
N
JA
FA
RER
AER
TBR
GPT-5.6-Sol
200
31.7
63.7
57.4
80.6
41.7
400
33.3
66.7
57.3
84.7
41.7
600
36.7
65.8
57.2
81.2
36.7
761 †
50.0
74.6
57.9
83.6
21.7
Gemini-3.6-Flash
200
0.0
7.5
51.1
42.9
90.0
400
0.0
5.2
49.8
39.6
93.3
Appendix
Table 9: Same 60 CoRR questions at increasing visible corpus size N . JA/FA/RER/AER/TBR in percent. † : reused main-batch endpoint.