Ramp Metering Control via Hybrid State Deep Reinforcement Learning in Partially Observable Connected Vehicle Environments
Organizations: Cosys-Grettia, Univ Gustave Eiffel, F-77454 Marne-la-Vallee, France. · Laboratoire LITAN, ESTIN, Amizour, Algeria.
Abstract
Freeway on-ramp merges are major sources of congestion, causing significant economic and environmental costs. While Deep Reinforcement Learning (DRL) offers a promising solution for ramp metering, existing approaches rely primarily on aggregated macroscopic data. Connected vehicles (CVs) provide vehicle-level observations that can complement aggregate traffic measurements, but their limited penetration produces incomplete microscopic information. This paper proposes a hybrid observation representation combining macroscopic traffic measurements with a two-channel grid encoding observed CV presence and speed. A Dueling Double Deep Q-Network processes these inputs to select ramp-metering green durations. The controller is trained under varying traffic demands and CV penetration rates and evaluated against ALINEA and macroscopic-only DRL variants in SUMO. Across 50 matched evaluation scenarios, the hybrid controller under partial CV visibility reduces the reported total travel time by 11.4 % and mean spillback duration by 84.9 % relative to ALINEA. Evaluating the same trained policy with full CV visibility yields a further travel-time reduction of approximately 1.6 %. Analysis across penetration rates suggests that the performance gap decreases as microscopic observations become more complete. These results support the use of complementary macroscopic and sparse microscopic observations for learning-based ramp metering. The source code implementation of the model is available at: https://github.com/youcefMehamlia/Multimodal-DRL-RMC
Figures & tables
| Action index | 0 | 1 | 2 | 3 | 4 | 5 | 6 | 7 |
| Green time (s) | 5 | 10 | 15 | 20 | 25 | 30 | 35 | 40 |
| Strategy | Travel Time (s) | Time Loss (s) | Wait Time (s) | Spillback (s) | CO 2 ( mg) |
|---|---|---|---|---|---|
| No Control | |||||
| ALINEA | |||||
| DQN Macro (No Lane) | |||||
| DQN Macro + Lane | |||||
| DQN Hybrid (Partial) | |||||
| DQN Hybrid (Full) |
Appendix figures & tables3 assets
Supplementary material from the paper’s appendix.
Appendix
| Edge ID | Description | Length (m) | Lanes | Speed (m/s) |
|---|---|---|---|---|
| main_road | Upstream mainline | 488.3 | 3 | 27.78 |
| on_ramp | Ramp to signal | 204.4 | 1 | 13.89 |
| passage_area | Signal to merge | 42.5 | 1 | 13.89 |
| accel_area | Merge / acc. lane | 193.8 | 4 | 22.22 |
| end_main | Downstream section | 193.1 | 3 | 27.78 |
| off_ramp | Diverging off-ramp | 161.3 | 2 | 13.89 |
| Component | Specification | Activation |
|---|---|---|
| Microscopic (CNN) | Conv2D (32, 64, 64 filters) | ELU |
| Macroscopic (FC) | Input: 14 features FC(128) | ELU |
| Shared Dense | FC(512) FC(256) | ELU |
| Value Head | FC(256) FC(1) | Linear |
| Advantage Head | FC(256) FC(8) | Linear |
| Hyperparameter | Value | Reward Weight | Value |
|---|---|---|---|
| Total training steps | (merge speed) | 1.5 | |
| Learning rate ( ) | (up speed) | 1.0 | |
| Discount factor ( ) | (down speed) | 0.5 | |
| Mini-batch size | 32 | (merge occ.) | 2.0 |
| Replay buffer size | (up occ.) | 1.0 | |
| Target update ( ) | (queue penalty) | 1.0 |