math.OCSep 28, 2026

Coordinated Lane-Level Variable Speed Limits and Ramp Metering for Successive Weaving Segments Considering Merging/Diverging Risks: A Hybrid Model Predictive Control and Multi-Agent Reinforcement Learning Approach

Authors: Guodong Ma, Baofeng Sun, Wenyu Yang, Zhihong Yao

Organizations: School of Transportation, Jilin University, Changchun 130022, China · School of Transportation and Logistics, Southwest Jiaotong University, Chengdu 610031, China

Abstract

Successive weaving segments (SWSs) on urban expressways are bottlenecks prone to recurrent congestion and collisions, requiring fine-grained active traffic management (ATM). Existing approaches struggle to balance the adaptive performance of data-driven optimization with the resilience and transferability of model-based control. We propose a hybrid framework to coordinate lane-level variable speed limits (VSLs) and ramp metering across SWSs. First, we reconstruct L-METANET, a lane-level macroscopic traffic flow model that captures free and forced lane changes. Second, we combine XGBoost-SHAP with a random-parameters binary logit (RPBL) model to derive analytical equations for merging and diverging collision risks and formulate system cost and reward functions. Third, we develop MPC-STMAPPO, a hierarchical controller integrating model predictive control (MPC) and multi-agent reinforcement learning (MARL). Its upper MPC layer uses L-METANET for long-horizon rolling optimization and generates baseline commands; its lower spatiotemporal MAPPO (ST-MAPPO) layer, enhanced with Mamba cells and graph attention, produces residual actions for short-horizon adjustment. Real-world experiments on the 18-km Eastern Expressway in Changchun, China, show that L-METANET accurately reproduces lane-changing-induced flow redistribution and capacity drops, with state evolution aligned with ground truth. XGBoost-SHAP-RPBL achieves AUCs above 0.80 in most tasks, outperforming conventional logit models. MPC-STMAPPO converges faster and performs better across multiple metrics than MPC- and MARL-based baselines. Under randomly fluctuating demand, it also significantly outperforms pure MARL in generalization, demonstrating strong potential for industrial deployment.

Explore similar work

CardsList
  1. HiLLTS: Zero-Shot Hierarchical LLM-Guided Traffic Signal Control for Sustainable Transportation

    Jul 17, 2026Yue Ding, Tendai Mukande, Mingming LiuTraffic Signal ControlCongestion

  2. Highway Congestion Reduction through Reinforcement Learning Based Eulerian Headway Control

    Dec 3, 2024Yaron Veksler, Sharon Hornstein, Han Wang +2Standard Traffic Speed BenchmarksHuman Driving

  3. Learning Agent-based Model Predictive Control for Holistic Vehicle Performance

    Sep 11, 2026Jiaming Zhong, Reza Valiollahi Mehrizi, Mohammad Pirani +4Model Predictive ControlGaussian Process