Effective traffic signal control (TSC) requires policies that respond to changing traffic demand and network conditions while meeting different control objectives. However, adapting existing strategies often involves repeated manual design and adjustment, making it difficult to systematically explore better control rules for a target network. Large language models (LLMs) can automate this process, but directly using them to select signal phases leaves decision rules embedded in black-box models and incurs recurring inference costs and latency. This paper formulates TSC as a modular program design problem and proposes EvoSignal, an LLM-guided evolutionary framework using traffic knowledge and performance feedback. The modular representation separates traffic feature extraction, local phase prioritization, and optional network-based priority adjustment. Starting from several established strategies, EvoSignal improves programs through feedback on congestion and signal operation, retaining strategies with different performance trade-offs. The resulting programs operate without online LLM inference. Simulation experiments across five scenarios on two real-world road networks show that the selected default EvoSignal program reduces waiting time by 16.8--49.2% relative to the lowest waiting time achieved by the 20 conventional, reinforcement learning-based, and LLM-based baselines in each scenario. A program prioritizing travel time and queue length outperforms all 20 baselines on all three metrics in the search scenario and remains among the top three on each metric when transferred unchanged to the other four scenarios. These findings support automated design of inspectable control programs that transfer across the evaluated road networks and traffic demands.Code is available at https://github.com/georgewanglz2019/EvoSignal.
Figures & tables
Figure 1: Modular control program at decision step t . One controller is instantiated per intersection, with controller i expanded; the selected phases jointly control the traffic environment.
Figure 2: EvoSignal framework for evolving modular control programs from multiple initial strategies, guided by traffic diagnostics and reusable strategy summaries. The selected program controls signals without online LLM inference.
Figure 3: Traffic networks, selectable phases, and demand profiles. (a,b) Jinan and Hangzhou layouts. (c) Four selectable phases. (d,e) Observed demand on a common scale. (f) J-high on a separate scale. Arrivals use five-minute intervals.
Demand scenario
ID
Signalized intersections
Hours
Vehicles
Arrivals (vehicles / 5 min)
Mean
SD
Min
Max
Core demand profiles
Jinan 1
J1
12
1
6,295
524.58
98.53
256
672
Jinan 2
J2
1
4,365
363.75
74.66
237
493
Jinan 3
J3
1
5,494
457.83
46.22
363
544
Jinan Extreme †
J-high
1
24,000
2000.00
44.21
1,933
2,098
Table 1: Traffic demand statistics for the two road networks.
Method
J1
J2
J3
TT (s)
QL (veh)
WT (s)
TT (s)
QL (veh)
WT (s)
TT (s)
QL (veh)
WT (s)
Traditional controllers
Random
597.62
687.35
99.46
555.23
428.38
100.40
552.74
529.63
99.33
FixedTime
481.79
491.03
70.99
441.19
294.14
66.72
450.11
394.34
69.19
Maxpressure
281.58
170.71
44.53
273.20
106.58
38.25
265.75
133.90
40.20
Reinforcement learning
Table 2: Traffic control performance on Jinan. Lower metric values and ranks are better.
Method
H1
H2
TT (s)
QL (veh)
WT (s)
TT (s)
QL (veh)
WT (s)
Traditional controllers
Random
621.14
295.81
96.06
504.28
432.92
92.61
FixedTime
616.02
301.33
73.99
486.72
425.15
72.80
Maxpressure
325.33
68.99
49.60
347.74
215.53
70.58
Reinforcement learning
Table 3: Traffic control performance on Hangzhou. Lower metric values and ranks are better.
Figure 4: Relative improvement over Maxpressure for two EvoSignal programs evolved on J1 and reused unchanged across five scenarios. Positive values indicate improvement.
Figure 5: Evolution-framework comparison on J1 under the default objective. (a) Mean best-so-far score with SD bands over three runs. (b) Distribution of final best scores.
Figure 6: Component ablations on J1 under the default objective: (a) best-so-far score over 200 evolution iterations; (b) final best scores. Each setting uses three runs, with the same plotting conventions as Figure 5 .
Figure 7: Objective-space distributions under the default and TT–QL objectives on J1: (a) travel time versus queue length; (b) queue length versus waiting time. Additional views and plotting details are provided in Appendix E .
Component
EvoSignal
EvoSignal (TT–QL)
M1: features
Pressure, waiting and approaching traffic, upstream occupancy, downstream load, and time since a phase was requested.
Pressure, waiting and approaching traffic, maximum downstream occupancy, and time since a phase was observed active.
M2: service
Service-age compensation grows with both age and waiting demand.
Service-age compensation is capped; the waiting weight adapts to congestion.
M2: congestion
Downstream load attenuates the pressure term; upstream occupancy amplifies demand.
High downstream occupancy attenuates the combined phase score.
M2: persistence
Applies different discounts to current and competing phase priorities at regular decision points.
Adds a continuation bonus when approaching demand remains relatively high.
M3: network adjustment
Favors the through phase with greater network directional pressure; adds a common offset based on neighboring intersection loads.
No additional adjustment.
Table 4: Control rules in the three modules of the two selected J1 programs. Details are provided in Appendix C .
Appendix figures & tables8 assets
Supplementary material from the paper’s appendix.
Appendix
Configuration
Runs
Mean ± SD
Best run
EvoSignal
3
0.397314 ± 0.000489
0.397742
OpenEvolve
3
0.394427 ± 0.002986
0.397484
ShinkaEvolve
3
0.395840 ± 0.001751
0.397672
Without SPA
3
0.394508 ± 0.002684
0.396541
Without PSA
3
0.396607 ± 0.000639
0.397309
Without 3M
3
0.393988 ± 0.003230
0.397682
Appendix
Table 5: Final best search scores on J1 under the default objective (Section 4.1.2 ).
Define the control task, editable modules, and current M2 focus.
Parent and references ( pa,Rn )
Supply executable code and alternative control examples.
Performance and SPA ( ma )
Report parent metrics, congestion patterns, and suggested changes.
PSA summary ( ba in ma )
Describe retained strategies and their measured trade-offs.
Output instructions
Request explanations followed by exact SEARCH/REPLACE edits.
Appendix
Table 8: Components of the recorded revision prompt.
Metric
Parent
Revised program
TT (s)
277.70
276.37
QL (veh)
163.89
161.73
WT (s)
36.52
35.64
Search score
0.392953
0.396055
Appendix
Table 9: Recorded parent and child performance for the illustrated revision.
Figure 8: Additional objective-space views of the default and TT–QL searches on J1: (a) three objectives; (b–d) pairwise projections over wider display ranges.
Setting
Value
Program evolution and LLM generation
Evolution attempts after initialization, N
200
Initial programs (full / without Multi-init)
11 / 1
LLM sampling probabilities (Flash / Pro)
0.7 / 0.3
Generation temperature / maximum output tokens
0.7 / 16,000
Population size / database archive size / islands
40 / 15 / 2
Appendix
Table 10: Core evolution and traffic-control settings.
Initial program
Design emphasis
Main control idea
Pressure–demand
Pressure and demand
Combine queue pressure, waiting vehicles, and approaching vehicles to rank phases.
Efficient pressure
Pressure-based control
Use efficient queue pressure directly as a simple initial priority rule.
Continuation-aware pressure
Phase continuation
Favor the current phase when approaching demand justifies continuation; otherwise follow pressure priorities.
Demand–supply balance
Demand and capacity
Combine estimated demand–supply ratios, saturation, and waiting vehicles to assess service urgency.
PID-inspired
Feedback control
Combine pressure, an accumulated-waiting proxy, and queue growth to respond to current and past congestion.
Fairness-aware
Service fairness
Increase priorities for phases not recently selected and reduce preference for a phase retained for a long period.
Appendix
Table 11: Initial control programs and their complementary design principles.
Urban traffic congestion significantly increases fuel consumption, greenhouse gas emissions, and commuter delays, resulting in substantial economic losses and environmental harm in modern cities. Traditional traffic signal control strategies such as fixed-time scheduling, actuated control, and reinforcement learning (RL)-based methods, offer different degrees of adaptability; however, RL-based methods can require extensive retraining, careful reward design, and substantial simulation data when transferred across networks or demand regimes. To address these challenges, we propose HiLLTS, an LLM-guided traffic signal control framework that employs a hierarchical three-layer architecture consisting of a central coordination agent, a district layer and multiple cluster-level intersection agents. Experimental results demonstrate consistent improvements in both congestion and environmental performance. Compared with the strongest non-LLM baseline in each scenario, HiLLTS reduces average waiting time by 36.73% under the low-congestion scenario and 14.71% under the high-congestion scenario, while reducing average CO2 emissions by 7.87% and 8.57%, respectively. Larger gains are observed against weaker baselines: under low congestion, HiLLTS achieves reductions of up to 18.00% in emissions and 62.07% in waiting time relative to Fixed-Time control; under high congestion, reductions of up to 28.89% in emissions and 40.36% in waiting time are observed relative to Max Pressure. The ablation study further validates the contribution of LLM-guided coordination over rule-based control
Yue Ding, Tendai Mukande, Mingming Liu
Research Ireland Centre for Researching Training in Machine Learning at Dublin City University
Traffic signal control (TSC) plays a central role in reducing congestion and maintaining urban mobility. This dissertation introduces DGLight, a critic-guided reinforcement-learning framework for adapting a pretrained large language model to TSC. DGLight first trains a CoLight-based Deep Q-Network critic to estimate traffic-aware action values from structured intersection states, then uses the frozen critic to score candidate language-model actions and optimize the policy with Group Relative Policy Optimization (GRPO). The resulting controller maps traffic states to interpretable reasoning traces and signal decisions while learning from dense per-state supervision rather than raw cumulative environment rewards. Experiments on TSC benchmarks covering Jinan and Hangzhou show that DGLight is the strongest overall method among the compared LLM-based controllers, remains competitive with strong RL baselines, and transfers well to city datasets not used to fit the critic. Qualitative examples further show that the model's generated reasoning is interpretable and aligned with the chosen signal phase. The project code is available [here](https://github.com/yyccbb/FYPLLMTSC).
Traffic signal control is a critical task in intelligent transportation systems, yet conventional fixed-time and rule-based methods often struggle to adapt to dynamic traffic demand and provide limited decision interpretability. This study proposes an LLM-augmented traffic signal control framework that integrates LSTM-based short-term traffic state prediction, predictive phase selection, structured large language model reasoning, and safety-constrained action filtering. The LSTM module forecasts future queue length, waiting time, vehicle count, and lane occupancy based on recent intersection-level observations. A predictive controller then generates candidate signal actions, while the LLM module evaluates these actions using structured traffic-state inputs and produces congestion diagnoses, phase adjustment recommendations, and natural-language explanations. To ensure operational reliability, all LLM-generated recommendations are validated by a safety filter before execution. Simulation-based experiments in SUMO compare the proposed method with fixed-time control, rule-based control, and an LSTM-based predictive baseline under balanced demand, directional peak demand, and sudden surge scenarios. The results indicate that the proposed framework improves traffic efficiency, especially under dynamic and non-recurrent traffic conditions, while maintaining zero constraint violations after safety filtering. Overall, this study demonstrates that LLMs can enhance traffic signal control when used as constrained reasoning and decision-support modules rather than direct low-level controllers. Keywords: Intelligent Transportation Systems; Traffic Signal Control; Large Language Models; LSTM; Traffic State Prediction; Decision Support; Safety-Constrained Control; SUMO Simulation.
Jiazhao Shi
Tandon School of Engineering, New York University, 6 MetroTech Center, Brooklyn, NY 11201, USA