ARGOS: Reinforcement Learning-Driven Multidimensional Elasticity for Service Orchestration in the Computing Continuum
Authors: Javier Mateos-Bravo, Sergio Laso, Juan Luis Herrera, Ilir Murturi, Pantelis Frangoudis, Schahram Dustdar
Organizations: Department of Computer Science and Telematics Engineering, University of Extremadura, Spain · Global Process and Product Improvement S.L., Spain · Department of Mechatronics, University of Prishtina, Kosova · Distributed Systems Group, TU Wien, Austria · ICREA, Barcelona, Spain
Data-intensive services in the Computing Continuum must balance analytics quality, resource usage, and cost across heterogeneous nodes with limited and uneven capacity. This balance becomes especially difficult when resource scaling reaches capacity limits, because changes in demand and cluster pressure must then be absorbed without violating client-defined quality ranges. Existing orchestrators mainly adapt resources, placements, or replicas, while analytics requirements such as coverage, sample, and freshness remain fixed. This article presents ARGOS, the Adaptive Reinforcement Learning-Driven Governance for Orchestrated Services, an end-to-end controller that formulates multidimensional elasticity as a per-request Markov decision process over analytics quality and cluster pressure, supported by capacity-aware admission. ARGOS is evaluated under controlled workloads and time-varying multi-tenant arrivals on a heterogeneous cluster. Across the controlled scenarios, the deep reinforcement learning policies consistently outperform the non-learning baselines and approach the independently tuned best-fixed reference. A separate live evaluation reports improvements over the static midpoint under realistic and saturated arrivals, with no recorded CPU or memory violations but remaining coverage violations. These results support deep reinforcement learning as an adaptive mechanism for multidimensional elasticity when resource scaling alone is insufficient.
Figures & tables
Figure 1 : (A) Coverage, (B) sample, and (C) freshness fixed after admission. Padlocks mark what is held fixed in each use case.
Study
Approach and control target
Evaluation setting
Reported comparison
Autonomic allocation Tesauro et al. [2006]
Hybrid RL, server allocation
Web-server resource management
Model-based allocation
DeepRM Mao et al. [2016]
Deep RL, cluster resources
Simulation
Resource-management heuristics
Decima Mao et al. [2019]
RL, dataflow scheduling
Data-processing cluster
Scheduling alternatives
Workflow autoscaling Garí et al. [2022]
Q-learning, workflow autoscaling
Cloud workflow evaluation
Autoscaling alternatives
Application placement Goudarzi et al. [2023]
Distributed deep RL, IoT placement
Simulation and testbed
Placement alternatives
Dynamic thresholds Rossi et al. [2023]
RL-tuned thresholds, application scaling
Simulation and prototype
Scaling policies
Table 1 : Representative orchestration approaches and evaluation designs.
Node
Role
vCPUs
Memory (GB)
Large
Orchestrator and worker
16
32
Medium
Worker
2
4
Small
Worker
1
2
Table 2 : Testbed nodes of the reference deployment.
Profile
Coverage
Sample
Freshness (s)
Lax-background
[0.25,0.45]
[0.20,0.45]
[90,180]
Short-burst
[0.60,0.95]
[0.60,0.90]
[10,45]
Aggressive-incident
[0.75,1.00]
[0.70,1.00]
[5,30]
Standard-operations
[0.45,0.75]
[0.35,0.65]
[30,90]
Cost-sensitive
[0.33,0.67]
[0.25,0.55]
[60,150]
Table 3 : Canonical workload profiles.
Runtime
Controller
vs. Static
vs. Threshold
vs. Best-fixed
Best-fixed recovered
Thread
DQN
+0.176 (5/5)
+0.392 (5/5)
−0.129 (0/5)
76%
PPO
+0.139 (5/5)
+0.355 (5/5)
−0.166 (0/5)
68%
Q-learning
−0.045 (1/5)
+0.171 (5/5)
−0.350 (0/5)
34%
Process
DQN
+0.179 (5/5)
+0.425 (5/5)
−0.142 (0/5)
72%
PPO
+0.151 (5/5)
+0.398 (5/5)
−0.170 (0/5)
67%
Q-learning
−0.043 (2/5)
+0.203 (5/5)
−0.364 (0/5)
29%
Table 4 : Paired reward advantage of each learned controller over each comparator.
Thread
Process
Controller
Fidelity
Duty
CPU p95
Fidelity
Duty
CPU p95
Best-fixed
0.950
16.8%
89%
0.950
16.2%
66%
DQN
0.940
13.8%
87%
0.939
18.1%
76%
PPO
0.931
10.9%
81%
0.933
12.7%
59%
Q-learning
0.919
12.3%
80%
0.917
8.5%
73%
Threshold
0.913
4.9%
58%
0.910
3.7%
51%
Table 5 : Spatial fidelity, service duty, and CPU utilization per controller; p95 denotes the 95th percentile.
Regime
Variant
Evaluated
Violated epochs
Coverage viol.
Freshness viol.
Sample/CPU/mem. viol.
Realistic
Static
51
39.0%
2335
0
0
DQN
50
71.1%
4246
0
0
PPO
50
60.5%
3684
1
0
Concurrency
Static
70
38.2%
4924
0
0
DQN
75
40.1%
5525
2
0
PPO
73
50.7%
6416
12
0
Table 6 : Live contract exposure per regime and variant.
A properly calibrated rule-based autoscaler can beat every one of six mainstream deep reinforcement learning (DRL) algorithms on cost across every workload we test - so when, if ever, does DRL actually help? We study this in RLScale-Bench, a reproducible benchmark and evaluation protocol for DRL on adaptive resource control, where an agent allocates compute to a dynamic workload under cost and service-level constraints. We evaluate PPO, DQN, A2C, SAC, TD3, and DDPG under matched architectures, training budgets, and reward functions against a calibrated rule-based baseline across six workload patterns and five seeds (240 runs), instantiate the benchmark on Kubernetes Horizontal Pod Autoscaling, and probe distribution-shift generalization. Three findings challenge common assumptions: (i) the calibrated controller achieves the lowest cost on all six workloads, though it trails the best RL agents on bursty and flash traffic; (ii) discrete-action algorithms outperform continuous-action ones by one to two orders of magnitude in constraint violations due to action-space mismatch; and (iii) no single algorithm dominates across workloads, with rankings shifting by up to four positions. The bottleneck in RL-based resource control is not algorithm selection but baseline calibration, reward engineering, and realistic evaluation protocols.
Guilin Zhang, Chuanyi Sun, Kai Zhao +3
The George Washington University, Washington, DC, USA
Reinforcement learning for service orchestration has been the subject of sustained research for over a decade, yet it is not used in production at scale. The usual explanation is that learned controllers degrade under delayed and noisy telemetry, workload shifts, and uncontrolled tenants. We test whether existing evidence supports that explanation. We evaluate three highly influential RL-based orchestration systems spanning resource allocation, DAG scheduling, and autoscaling, using pre-registered predictions about comparative degradation under production-relevant perturbations and paired inference with family-wise error correction. Across the tests, most predicted performance reversals do not occur. Diagnostic analyses show that these outcomes often reflect comparator collapse, artefact limitations, or evaluation choices rather than evidence that learned controllers tolerate the perturbations. One apparent advantage under observation lag is roughly fortyfold compared to a Kubernetes HPA-equivalent controller. Another widely cited result cannot be reconstructed from its released artefact, and the strongest reproducible margin is far smaller than the published results. Conclusions also reverse under changes in perturbation magnitude and evaluation mode. Based on these results and broader patterns in the literature, we identify an institutional problem. Publication and review incentives favour benchmark gains against convenient comparators, even when those gains provide little evidence of deployment performance. We argue that the problem is not solely technical. Rather, it is institutional, so learned orchestration needs production-grade comparators, registered perturbation models, separate operational metrics, and publication criteria that reward reproducible operational evidence. Without these changes, the literature can grow without establishing whether learning improves orchestration.
In edge computing, the stochastic and bursty nature of serverless workloads challenges autonomous resource orchestration. Traditional reactive controllers, such as the Kubernetes Horizontal Pod Autoscaler (HPA), suffer from reaction latency, leading to Service Level Objective (SLO) violations during traffic spikes and resource flapping during ramp-downs. While Deep Reinforcement Learning (DRL) offers a pathway toward proactive management, standard agents suffer from \textit{temporal blindness}, an inability to exploit the recent temporal context in non-Markovian edge environments. To bridge this gap, we propose a stability-aware autoscaling framework unifying short-horizon temporal context and control via an Attention-Enhanced Double-Stacked LSTM architecture integrated within a Proximal Policy Optimization (PPO) agent. Unlike shallow recurrent models, our approach employs a learned attention mechanism that weights recent historical states non-uniformly, suppressing high-frequency jitter while preserving the trend that precedes demand shifts. We validate the framework on two independent Kubernetes clusters using real-world Azure Functions traces. Against the single-layer LSTM ablation and the static HPA baseline, our approach reduces P90 latency by ≈67%, and holds average latency within the 50ms hard SLO for 98.8% of the run against 49.6% and 43.5% respectively. Against Kubernetes Event-Driven Autoscaling (KEDA), it matches latency performance at 75% fewer replica-steps and 59% less churn, with P90 hard-SLO violation bursts of at most 5 consecutive intervals against up to 24 for KEDA. These results indicate that mitigating temporal blindness through deep attentive memory improves the reliability and stability of Kubernetes autoscaling under bursty edge workloads.
Faraz Shaikh, Gianluca Reali, Mauro Femminella
Department of Engineering, University of Perugia, Perugia, Italy · Consorzio Nazionale Interuniversitario per le Telecomunicazioni (CNIT), 43124 Parma, Italy