VLALight: A Vision-Language-Action Model for Traffic Signal Control
Organizations: The Hong Kong University of Science and Technology (Guangzhou) · Dalian University of Technology
Abstract
Traffic signal control (TSC) is essential for improving urban mobility and reducing congestion. Although roadside cameras are widely deployed at signalized intersections and provide rich visual observations of evolving traffic, existing TSC methods typically rely on manually engineered traffic states or separate perception modules, creating a gap between physical observations and control decisions. We present VLALight, the first vision-language-action (VLA) model for end-to-end traffic signal control from multi-view roadside videos. VLALight directly maps visual observations to coordinated signal actions through multi-target spatiotemporal traffic reasoning and topology-aware cooperative perception across intersections. To establish this capability, we develop a two-stage supervised cold-start training strategy for visual traffic understanding and signal decision-making, followed by cooperative agentic reinforcement learning that jointly optimizes local control and network-wide traffic efficiency. Furthermore, VLALight introduces adaptive fast and slow reasoning modes, enabling the policy to allocate deeper reasoning only when additional deliberation provides sufficient control benefits. Through balanced mode-aware rollouts and relative advantage optimization, VLALight learns to trade off decision quality and inference cost. Extensive experiments on seven real-world traffic-flow datasets across three urban networks demonstrate that VLALight consistently outperforms transportation-based, RL-based, and LLM/VLM-based baselines. Ablation studies validate the effectiveness of cooperative perception, network-level optimization, and adaptive reasoning. These results demonstrate the potential of VLA models for real-world physical traffic control. Our project is available at https://github.com/usail-hkust/VLALight.git.
Figures & tables
| Method | Jinan | Hangzhou | New York | ||||||||||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| 1 | 2 | 3 | 1 | 2 | 1 | 2 | |||||||||||||||
| ATT | AQL | AWT | ATT | AQL | AWT | ATT | AQL | AWT | ATT | AQL | AWT | ATT | AQL | AWT | ATT | AQL | AWT | ATT | AQL | AWT | |
| Transportation-Engineering Methods | |||||||||||||||||||||
| FixedTime | 458.95 | 383.91 | 237.36 | 366.59 | 185.60 | 152.92 | 394.99 | 206.57 | 136.14 | 544.46 | 207.24 | 269.66 | 464.08 | 274.97 | 229.64 | 1464.24 | 2716.68 | 1171.09 | 1666.71 | 3940.70 | 1387.68 |
| MaxPressure | 270.97 | 163.77 | 79.79 | 267.68 | 109.82 | 74.35 | 260.02 | 135.30 | 73.35 | 306.45 | 65.22 | 63.13 | 304.98 | 120.88 | 73.96 | 1215.43 | 2395.16 | 894.25 | 1467.47 | 4066.03 | 1202.50 |
| RL-Based Methods | |||||||||||||||||||||
| Trace | Method | Reasoning statistics | Control performance | |||
|---|---|---|---|---|---|---|
| Fast (%) | Slow (%) | Avg. Tokens (Slow) | Min–max Tokens (Slow) | ATT | ||
| New York 1 | VLALight | 51.92 | 48.08 | 220.47 | 54–545 | 975.80 |
| New York 1 | w/o balanced rollouts | 27.56 | 72.44 | 209.42 | 62–649 | 984.76 |
| New York 1 | SFT | 59.72 | 40.28 | 233.15 | 62–1,230 | 1171.08 |
| New York 2 | VLALight | 51.24 | 48.76 | 208.26 | 58–682 | 1240.11 |
| New York 2 | w/o balanced rollouts | 28.38 | 71.62 | 204.92 | 62–609 | 1296.10 |
| Fast 1 clear local winner intersection 3–1, step 99 Local perception ETWT: =19, =19, =0, =0 NTST : = 39 , = 39 , = 0 , = 0 ELWL: =8, =8, =0, =0 NLSL: =0, =0, =0, =0 Cooperative context Neighbor movement cues: ETWT: 19 vehicles NTST: 19 vehicles Neighbor directional pressure: N: 20 queued E: 19 queued ETA: 27 s FAST signal NTST Why: The local queue maximum is unambiguous. | Fast 2 clear local winner intersection 5–1, step 88 Local perception ETWT: =14, =0, = , = NTST : = 21 , = 21 , = 0 , = 0 ELWL: =1, =1, =0, =0 NLSL: =0, =0, =0, =0 Cooperative context Neighbor movement cues: ETWT: 1 vehicle Neighbor directional pressure: E: 1 queued ETA: 27 s FAST signal NTST Why: One phase dominates visible and stopped vehicles. |
| Slow 1 competing local queues intersection 2–1, step 40 Local perception ETWT : = 20 , = 19 , = , = NTST: =21, =19, = , = ELWL: =5, =5, =0, =0 NLSL: =0, =0, =0, =0 Cooperative context Neighbor movement cues: ETWT: 11 vehicles NTST: 1 vehicle Neighbor directional pressure: E: 20 queued N: 12 queued ETA: 27 s SLOW signal ETWT Key reasoning: Given ETWT has the largest visible demand, worsening trend, and strong local coordination plus east neighbor pressure, it is clearly the most effective choice. | Slow 2 close local competition intersection 7–10, step 62 Local perception ETWT: =4, =3, = , = NTST: =0, =0, = , = ELWL : = 5 , = 3 , = , = NLSL: =2, =2, =0, = Cooperative context Neighbor movement cues: ETWT: 1 vehicle ELWL: 1 vehicle Neighbor directional pressure: N: 2 queued S: 2 queued ETA: 27 s SLOW signal ELWL Key reasoning: Given ELWL’s high demand, queue, positive trend, and long waiting time, it is the most effective choice. |
Appendix figures & tables5 assets
Supplementary material from the paper’s appendix.
Appendix
| City | Trace | Grid | Vehicles | 5-min arrivals ( ) | Min–max |
|---|---|---|---|---|---|
| Jinan | 1 | 6295 | 256–672 | ||
| Jinan | 2 | 4365 | 237–493 | ||
| Jinan | 3 | 5494 | 363–544 | ||
| Hangzhou | 1 | 2983 | 212–333 | ||
| Hangzhou | 2 | 6984 | 203–1146 | ||
| New York | 1 | 11058 | 383–965 |