Efficient resource provisioning for large-scale workflows on cloud infrastructures is a critical performance engineering challenge. These workflows are often structured as directed acyclic graphs (DAGs), where under-provisioning can cause critical bottlenecks and over-provisioning leads to unnecessary costs. Accurate, task-level prediction of resource intensity (e.g., CPU load and memory usage) is essential for mitigating these issues. While task-level features are commonly used for prediction, the performance impact of the workflow's overall topological structure is often overlooked or assumed. The central question of our work is: To what extent does what part of the DAG topology influence task-level resource intensity, and what is the most effective way to model this influence? This paper presents a comprehensive benchmark to systematically quantify the impact of graph topology on task intensity prediction. We evaluate and compare a spectrum of modeling approaches. Our findings demonstrate that topology is a critical feature for accurate prediction. Models incorporating important topological information, even through simple handcrafted features, significantly outperform baseline models. We show that graph-native models provide the highest accuracy, achieving low mean absolute errors for both CPU and memory predictions, and can still be combined with simple topological features that they do not learn for better performance.
Figures & tables
Fig. 1 : Example workflow DAG with CPU/memory intensities.
Feature Method
CPU Accuracy (1 st Split)
CPU Accuracy (2 nd Split)
CPU Accuracy (3 rd Split)
Memory Accuracy
DAG In & Out
51.87%
52.18%
51.99%
91.65%
Node Degrees
48.27%
48.24%
48.15%
90.98%
Clustering Coefficient
48.14%
48.25%
48.18%
90.98%
Graphlets
48.27%
48.10%
48.39%
90.98%
Node 1 Path Length
70.74%
70.50%
70.00%
90.98%
Node 2 Path Length
65.55%
63.30%
64.58%
90.98%
TABLE I : Impact of topological features on CPU and memory intensity prediction (experiment 1).
Feature Method
CPU Accuracy (1 st Split)
CPU Accuracy (2 nd Split)
CPU Accuracy (3 rd Split)
Memory Accuracy
Only Measured Features
86.60%
86.82%
87.00%
96.21%
Only All Topology Features
77.63%
77.20%
77.55%
95.04%
DAG In & Out
90.16%
89.43%
90.22%
98.51%
Clustering Coefficient
88.59%
87.53%
87.83%
97.80%
Node Degrees
88.33%
87.40%
88.86%
97.72%
Graphlets
87.98%
87.66%
89.25%
97.76%
TABLE II : Impact of topological features when topology information is considered for all nodes (experiment 2).
GNN
CPU Acc. (1 st Split)
CPU Acc. (2 nd Split)
CPU Acc. (3 rd Split)
GCN
88.38±0.4%
88.18±0.2%
88.91±0.1%
GAT
89.81±0.2%
89.61±0.2%
90.48±0.3%
GIN
89.87±0.1%
89.63±0.2%
90.64±0.3%
GraphSAGE
87.80±0.2%
87.35±0.2%
88.18±0.2%
H2GCN
87.84±0.1%
87.11±0.1%
88.20±0.1%
DAGTransformer (stated in paper)
91.25±0.04%
91.11±0.05%
92.15±0.13%
TABLE III : Graph learning models CPU average prediction accuracies without topological features.
GNN
CPU Acc. (1 st Split)
CPU Acc. (2 nd Split)
CPU Acc. (3 rd Split)
GCN
90.34±0.3%
89.67±0.1%
90.35±0.4%
GAT
90.54±0.1%
89.89±0.1%
90.89±0.2%
GIN
90.56±0.1%
90.14±0.2%
90.74±0.05%
GraphSAGE
88.97±0.4%
88.08±0.3%
89.10±0.1%
H2GCN
89.31±0.1%
88.73±0.3%
89.81±0.2%
DAGTransformer (stated in paper)
-
-
-
TABLE IV : Graph learning models CPU average prediction accuracies with topological features.
GNN
Memory Accuracy
Memory Accuracy with Topological Features
GCN
97.67±0.03%
98.52±0.03%
GAT
98.34±0.1%
98.68±0.1%
GIN
98.47±0.2%
98.59±0.1%
GraphSAGE
97.33±0.1%
97.78±0.4%
H2GCN
90.98±0.01%
90.98±0.004%
DAGTransformer (stated)
98.56%
-
TABLE V : Graph learning models memory average prediction.
task
avg CPU
max CPU
avg mem
max mem
20 Mean
8.42
22.76
0.085
0.10
20 Mean Topo
7.77
22.51
0.074
0.09
20 Max
9.79
34.60
0.11
0.12
20 Max Topo
9.53
33.11
0.10
0.10
18 Mean
11.51
30.24
0.06
0.08
18 Mean Topo
10.67
28.42
0.06
0.08
TABLE VI : MAE values for average CPU & memory utilization.
task
avg CPU
max CPU
avg mem
max mem
20 Mean
0.85
0.41
0.95
0.94
20 Mean Topo
0.87
0.51
0.96
0.95
20 Max
0.83
0.45
0.93
0.92
20 Max Topo
0.84
0.50
0.94
0.94
18 Mean
0.79
0.68
0.41
0.61
18 Mean Topo
0.81
0.75
0.46
0.60
TABLE VII : R2 scores for average CPU & memory utilization.
Fig. 2 : Regression plots when model embeddings are used to feed XGBoost.
Dynamic cloud workflow scheduling must balance deadline satisfaction, container utilization, and energy consumption while dealing with stochastic task-execution speeds, placement-dependent communication, and coupled task and container decisions. Workflows are naturally modeled as directed acyclic graphs (DAGs), but conventional vector- or matrix-based states do not fully capture their dependency topology. To better represent task urgency and structural relationships, we assign predicted sub-deadlines to tasks and use a multi-head graph attention network (GAT) to extract dependency information from the evolving DAGs. Based on these representations, we develop a Graph Attention-Driven Hierarchical Reinforcement Learning (GA-HRL) framework and model the scheduling process as an event-driven hierarchical semi-Markov decision process (SMDP). Workflow arrivals and task completions trigger scheduling events. At each scheduling event, the Task Scheduling (TS) agent first processes the currently ready tasks by assigning them to admissible existing containers or requesting new ones. The requested containers are then processed by the Container Scheduling (CS) agent for host placement before the environment advances. The two agents are trained alternately using separate Proximal Policy Optimization (PPO). Experiments on the 2018 Alibaba cluster trace show that GA-HRL maintains competitive workflow success rate and, in settings where success is comparable, generally achieves higher container utilization and lower energy consumption. Under the largest speed variation, it trades a small success-rate margin for substantially lower energy. Simulation code is available at: https://github.com/zongjin130/GA-HRL.
Zongjin Li, Shaohan Feng, Chunxi Yang +1
Faculty of Mechanical and Electrical Engineering, Kunming University of Science and Technology, Kunming 650500, China · School of Information and Electronic Engineering (Sussex Artificial Intelligence Institute), Zhejiang Gongshang University, Hangzhou 310018, China
With the rapid development of the Internet of Things, computation intensive directed acyclic graph (DAG) tasks have become increasingly common in cloud-edge-end collaborative environments. However, cloud, edge, and end nodes are highly heterogeneous in computing capacity, network bandwidth, and energy consumption, which makes the efficient scheduling of tasks with complex dependencies an NP-hard problem. Traditional heuristic algorithms and conventional reinforcement-learning methods often fail to capture the spatio-temporal dynamics of system resources. This paper proposes PPO-STGNN, a DAG task-scheduling algorithm that integrates proximal policy optimization (PPO) with spatio-temporal graph neural networks (STGNNs). The method uses an STGNN to extract features from both the DAG task topology and the physical cloud-edge-end resource graph, and then optimizes the scheduling policy through PPO to minimize makespan and schedule length ratio (SLR) while improving CPU and memory load balancing. To accelerate convergence, a multi-teacher behavior-cloning mechanism is introduced for pretraining. Experimental results show that PPO-STGNN significantly improves load balancing while maintaining a low completion time, making it suitable for dynamic and heterogeneous cloud-edge- end DAG scheduling scenarios.
Yangshuo Qi, Chenwei Wang, Zihan Shen +1
Beijing University of Posts and Telecommunications, Beijing, China · Google, Mountain View, USA · NanJing University, NanJing, China
Accurate cloud workload forecasting is pivotal for efficient resource management but remains challenging as workloads are highly volatile and prone to sudden bursts. Although wavelets preserve temporal locality, rigid fixed bases struggle with complex patterns and isolated processing neglects critical spatial dependencies. To address this, we propose SWIFT, a pure convolutional framework designed for high-efficiency workload forecasting. We introduce a Learnable Cascaded Wavelet Path that reformulates the traditional fixed wavelet bases into adaptive convolutional operators, enabling precise, data-driven feature peeling. Complementing this, our Multivariate Interaction Module sequentially models inter-variable spatial and intra-variable feature interactions to stabilize and refine noisy workload states. Extensive experiments demonstrate that SWIFT achieves SOTA accuracy with linear O(L) complexity, reducing prediction error by up to 31.04% while cutting latency by 79.74%.
Zeyuan Ding, Lingfeng Zheng, Dian Ding +1
1Shanghai Jiao Tong University, Shanghai, China · 2Shaanxi Normal University, Xi’an, China