ARGOS: Reinforcement Learning-Driven Multidimensional Elasticity for Service Orchestration in the Computing Continuum
Authors: Javier Mateos-Bravo, Sergio Laso, Juan Luis Herrera, Ilir Murturi, Pantelis Frangoudis, Schahram Dustdar
Organizations: Department of Computer Science and Telematics Engineering, University of Extremadura, Spain · Global Process and Product Improvement S.L., Spain · Department of Mechatronics, University of Prishtina, Kosova · Distributed Systems Group, TU Wien, Austria · ICREA, Barcelona, Spain
Data-intensive services in the Computing Continuum must balance analytics quality, resource usage, and cost across heterogeneous nodes with limited and uneven capacity. This balance becomes especially difficult when resource scaling reaches capacity limits, because changes in demand and cluster pressure must then be absorbed without violating client-defined quality ranges. Existing orchestrators mainly adapt resources, placements, or replicas, while analytics requirements such as coverage, sample, and freshness remain fixed. This article presents ARGOS, the Adaptive Reinforcement Learning-Driven Governance for Orchestrated Services, an end-to-end controller that formulates multidimensional elasticity as a per-request Markov decision process over analytics quality and cluster pressure, supported by capacity-aware admission. ARGOS is evaluated under controlled workloads and time-varying multi-tenant arrivals on a heterogeneous cluster. Across the controlled scenarios, the deep reinforcement learning policies consistently outperform the non-learning baselines and approach the independently tuned best-fixed reference. A separate live evaluation reports improvements over the static midpoint under realistic and saturated arrivals, with no recorded CPU or memory violations but remaining coverage violations. These results support deep reinforcement learning as an adaptive mechanism for multidimensional elasticity when resource scaling alone is insufficient.
Figures & tables
Figure 1 : (A) Coverage, (B) sample, and (C) freshness fixed after admission. Padlocks mark what is held fixed in each use case.
Study
Approach and control target
Evaluation setting
Reported comparison
Autonomic allocation Tesauro et al. [2006]
Hybrid RL, server allocation
Web-server resource management
Model-based allocation
DeepRM Mao et al. [2016]
Deep RL, cluster resources
Simulation
Resource-management heuristics
Decima Mao et al. [2019]
RL, dataflow scheduling
Data-processing cluster
Scheduling alternatives
Workflow autoscaling Garí et al. [2022]
Q-learning, workflow autoscaling
Cloud workflow evaluation
Autoscaling alternatives
Application placement Goudarzi et al. [2023]
Distributed deep RL, IoT placement
Simulation and testbed
Placement alternatives
Dynamic thresholds Rossi et al. [2023]
RL-tuned thresholds, application scaling
Simulation and prototype
Scaling policies
Table 1 : Representative orchestration approaches and evaluation designs.
Node
Role
vCPUs
Memory (GB)
Large
Orchestrator and worker
16
32
Medium
Worker
2
4
Small
Worker
1
2
Table 2 : Testbed nodes of the reference deployment.
Profile
Coverage
Sample
Freshness (s)
Lax-background
[0.25,0.45]
[0.20,0.45]
[90,180]
Short-burst
[0.60,0.95]
[0.60,0.90]
[10,45]
Aggressive-incident
[0.75,1.00]
[0.70,1.00]
[5,30]
Standard-operations
[0.45,0.75]
[0.35,0.65]
[30,90]
Cost-sensitive
[0.33,0.67]
[0.25,0.55]
[60,150]
Table 3 : Canonical workload profiles.
Runtime
Controller
vs. Static
vs. Threshold
vs. Best-fixed
Best-fixed recovered
Thread
DQN
+0.176 (5/5)
+0.392 (5/5)
−0.129 (0/5)
76%
PPO
+0.139 (5/5)
+0.355 (5/5)
−0.166 (0/5)
68%
Q-learning
−0.045 (1/5)
+0.171 (5/5)
−0.350 (0/5)
34%
Process
DQN
+0.179 (5/5)
+0.425 (5/5)
−0.142 (0/5)
72%
PPO
+0.151 (5/5)
+0.398 (5/5)
−0.170 (0/5)
67%
Q-learning
−0.043 (2/5)
+0.203 (5/5)
−0.364 (0/5)
29%
Table 4 : Paired reward advantage of each learned controller over each comparator.
Thread
Process
Controller
Fidelity
Duty
CPU p95
Fidelity
Duty
CPU p95
Best-fixed
0.950
16.8%
89%
0.950
16.2%
66%
DQN
0.940
13.8%
87%
0.939
18.1%
76%
PPO
0.931
10.9%
81%
0.933
12.7%
59%
Q-learning
0.919
12.3%
80%
0.917
8.5%
73%
Threshold
0.913
4.9%
58%
0.910
3.7%
51%
Table 5 : Spatial fidelity, service duty, and CPU utilization per controller; p95 denotes the 95th percentile.
Regime
Variant
Evaluated
Violated epochs
Coverage viol.
Freshness viol.
Sample/CPU/mem. viol.
Realistic
Static
51
39.0%
2335
0
0
DQN
50
71.1%
4246
0
0
PPO
50
60.5%
3684
1
0
Concurrency
Static
70
38.2%
4924
0
0
DQN
75
40.1%
5525
2
0
PPO
73
50.7%
6416
12
0
Table 6 : Live contract exposure per regime and variant.
Department of Engineering, University of Perugia, Perugia, Italy · Consorzio Nazionale Interuniversitario per le Telecomunicazioni (CNIT), 43124 Parma, Italy