Deadline-Aware Multi-Agent Reinforcement Learning for TSN-Based Vehicular Edge Networks
Organizations: Universidade Federal de Minas Gerais, Brazil · School of Electrical Engineering and Computer Science, University of Ottawa, Ottawa, Canada · University of Antwerp - imec, IDLab - Faculty of Applied Engineering, Belgium
Abstract
Vehicular edge computing (VEC) enables latency-sensitive applications by bringing computing and networking resources closer to vehicles. However, existing approaches often overlook network contention among co-located services with heterogeneous and dynamic latency requirements. While time-sensitive networking (TSN) provides bounded-latency communication, conventional and reinforcement learning-based schedulers struggle to adapt to highly dynamic vehicular environments and inter-queue dependencies. To address these limitations, we propose a multi-agent reinforcement learning (MARL) approach for queue-level scheduling in TSN-enabled VEC. Each TSN queue is assigned an autonomous agent that jointly learns the queue service order and time-slot duration to minimize deadline misses under speed-dependent latency requirements. We employ multi-agent proximal policy optimization (MAPPO) to enable coordinated yet autonomous scheduling decisions. Evaluation against single-agent, multi-agent, and non-learning-based baselines shows that MAPPO provides robust performance across different traffic profiles. Compared with centralized single-agent methods, it reduces service latency by up to 66.2% and improves reliability by up to 271.8%. Furthermore, unlike urgency-based heuristics, MAPPO ensures balanced scheduling while achieving lower inference times compared to other MARL methods.
Figures & tables
| VEC | Value | MAPPO parameter | Value |
| Network speed (Gbps) | 10 | Learning rate | |
| Max packet size (B) | 1428 | gamma | 0.99 |
| Number of vehicles | 1–100 | Episode length | 150 |
| Average speed (km/h) | 10–120 | PPO epochs | 5 |
| AR setting/Base deadline | SeAR setting/Base deadline | ||
| q0, q3: 3840 1920@90 Hz / 11.11 ms | q2: 1920 720@60 Hz / 2 ms | ||
| Method | Inference Time | Method | Inference Time |
| MAPPO | 0.69 0.18 | MAT | 10.90 1.29 |
| HAPPO | 5.48 0.61 | PPO | 0.78 0.22 |
| A2C | 0.79 0.27 |