EvoSignal: LLM-Guided Evolutionary Design of Modular Traffic Signal Control Programs
Organizations: Monash University · KTH Royal Institute of Technology · Chinese Academy of Sciences · Zhejiang Dahua Technology · Southeast University
Abstract
Effective traffic signal control (TSC) requires policies that respond to changing traffic demand and network conditions while meeting different control objectives. However, adapting existing strategies often involves repeated manual design and adjustment, making it difficult to systematically explore better control rules for a target network. Large language models (LLMs) can automate this process, but directly using them to select signal phases leaves decision rules embedded in black-box models and incurs recurring inference costs and latency. This paper formulates TSC as a modular program design problem and proposes EvoSignal, an LLM-guided evolutionary framework using traffic knowledge and performance feedback. The modular representation separates traffic feature extraction, local phase prioritization, and optional network-based priority adjustment. Starting from several established strategies, EvoSignal improves programs through feedback on congestion and signal operation, retaining strategies with different performance trade-offs. The resulting programs operate without online LLM inference. Simulation experiments across five scenarios on two real-world road networks show that the selected default EvoSignal program reduces waiting time by 16.8--49.2% relative to the lowest waiting time achieved by the 20 conventional, reinforcement learning-based, and LLM-based baselines in each scenario. A program prioritizing travel time and queue length outperforms all 20 baselines on all three metrics in the search scenario and remains among the top three on each metric when transferred unchanged to the other four scenarios. These findings support automated design of inspectable control programs that transfer across the evaluated road networks and traffic demands.Code is available at https://github.com/georgewanglz2019/EvoSignal.
Figures & tables
| Demand scenario | ID | Signalized intersections | Hours | Vehicles | Arrivals (vehicles / 5 min) | |||
| Mean | SD | Min | Max | |||||
| Core demand profiles | ||||||||
| Jinan 1 | J1 | 12 | 1 | 6,295 | 524.58 | 98.53 | 256 | 672 |
| Jinan 2 | J2 | 1 | 4,365 | 363.75 | 74.66 | 237 | 493 | |
| Jinan 3 | J3 | 1 | 5,494 | 457.83 | 46.22 | 363 | 544 | |
| Jinan Extreme † | J-high | 1 | 24,000 | 2000.00 | 44.21 | 1,933 | 2,098 | |
| Method | J1 | J2 | J3 | ||||||
|---|---|---|---|---|---|---|---|---|---|
| TT (s) | QL (veh) | WT (s) | TT (s) | QL (veh) | WT (s) | TT (s) | QL (veh) | WT (s) | |
| Traditional controllers | |||||||||
| Random | 597.62 | 687.35 | 99.46 | 555.23 | 428.38 | 100.40 | 552.74 | 529.63 | 99.33 |
| FixedTime | 481.79 | 491.03 | 70.99 | 441.19 | 294.14 | 66.72 | 450.11 | 394.34 | 69.19 |
| Maxpressure | 281.58 | 170.71 | 44.53 | 273.20 | 106.58 | 38.25 | 265.75 | 133.90 | 40.20 |
| Reinforcement learning | |||||||||
| Method | H1 | H2 | ||||
|---|---|---|---|---|---|---|
| TT (s) | QL (veh) | WT (s) | TT (s) | QL (veh) | WT (s) | |
| Traditional controllers | ||||||
| Random | 621.14 | 295.81 | 96.06 | 504.28 | 432.92 | 92.61 |
| FixedTime | 616.02 | 301.33 | 73.99 | 486.72 | 425.15 | 72.80 |
| Maxpressure | 325.33 | 68.99 | 49.60 | 347.74 | 215.53 | 70.58 |
| Reinforcement learning | ||||||
| Component | EvoSignal | EvoSignal (TT–QL) |
|---|---|---|
| M1: features | Pressure, waiting and approaching traffic, upstream occupancy, downstream load, and time since a phase was requested. | Pressure, waiting and approaching traffic, maximum downstream occupancy, and time since a phase was observed active. |
| M2: service | Service-age compensation grows with both age and waiting demand. | Service-age compensation is capped; the waiting weight adapts to congestion. |
| M2: congestion | Downstream load attenuates the pressure term; upstream occupancy amplifies demand. | High downstream occupancy attenuates the combined phase score. |
| M2: persistence | Applies different discounts to current and competing phase priorities at regular decision points. | Adds a continuation bonus when approaching demand remains relatively high. |
| M3: network adjustment | Favors the through phase with greater network directional pressure; adds a common offset based on neighboring intersection loads. | No additional adjustment. |
Appendix figures & tables8 assets
Supplementary material from the paper’s appendix.
Appendix
| Configuration | Runs | Mean SD | Best run |
|---|---|---|---|
| EvoSignal | 3 | 0.397314 0.000489 | 0.397742 |
| OpenEvolve | 3 | 0.394427 0.002986 | 0.397484 |
| ShinkaEvolve | 3 | 0.395840 0.001751 | 0.397672 |
| Without SPA | 3 | 0.394508 0.002684 | 0.396541 |
| Without PSA | 3 | 0.396607 0.000639 | 0.397309 |
| Without 3M | 3 | 0.393988 0.003230 | 0.397682 |
| Scenario | Programs | TT (s) | QL (veh) | WT (s) |
|---|---|---|---|---|
| EvoSignal | ||||
| J1 | 3 | 275.97 1.86 | 160.64 2.50 | 35.64 1.52 |
| J2 | 3 | 268.21 0.51 | 101.75 0.50 | 33.06 1.43 |
| J3 | 3 | 261.57 1.00 | 128.56 0.56 | 34.34 0.96 |
| H1 | 3 | 316.15 2.18 | 61.76 1.32 | 29.88 0.47 |
| H2 | 3 | 329.69 1.62 | 177.71 2.69 | 34.45 4.92 |
| Scenario | Programs | TT (s) | QL (veh) | WT (s) |
|---|---|---|---|---|
| EvoSignal (H1, TT–QL) | ||||
| J1 | 3 | 276.66 3.74 | 162.54 4.26 | 40.44 1.67 |
| J2 | 3 | 268.40 0.26 | 102.08 0.00 | 37.29 1.36 |
| J3 | 3 | 260.31 1.06 | 127.36 1.43 | 38.97 2.06 |
| H1 | 3 | 307.77 0.59 | 55.86 0.59 | 32.34 2.17 |
| H2 | 3 | 332.12 3.71 | 186.25 6.36 | 52.92 2.34 |
| Input component | Role in program revision |
|---|---|
| Task and interfaces ( ) | Define the control task, editable modules, and current M2 focus. |
| Parent and references ( ) | Supply executable code and alternative control examples. |
| Performance and SPA ( ) | Report parent metrics, congestion patterns, and suggested changes. |
| PSA summary ( in ) | Describe retained strategies and their measured trade-offs. |
| Output instructions | Request explanations followed by exact SEARCH/REPLACE edits. |
| Metric | Parent | Revised program |
|---|---|---|
| TT (s) | 277.70 | 276.37 |
| QL (veh) | 163.89 | 161.73 |
| WT (s) | 36.52 | 35.64 |
| Search score | 0.392953 | 0.396055 |
| Setting | Value |
|---|---|
| Program evolution and LLM generation | |
| Evolution attempts after initialization, | 200 |
| Initial programs (full / without Multi-init) | 11 / 1 |
| LLM sampling probabilities (Flash / Pro) | 0.7 / 0.3 |
| Generation temperature / maximum output tokens | 0.7 / 16,000 |
| Population size / database archive size / islands | 40 / 15 / 2 |
| Initial program | Design emphasis | Main control idea |
|---|---|---|
| Pressure–demand | Pressure and demand | Combine queue pressure, waiting vehicles, and approaching vehicles to rank phases. |
| Efficient pressure | Pressure-based control | Use efficient queue pressure directly as a simple initial priority rule. |
| Continuation-aware pressure | Phase continuation | Favor the current phase when approaching demand justifies continuation; otherwise follow pressure priorities. |
| Demand–supply balance | Demand and capacity | Combine estimated demand–supply ratios, saturation, and waiting vehicles to assess service urgency. |
| PID-inspired | Feedback control | Combine pressure, an accumulated-waiting proxy, and queue growth to respond to current and past congestion. |
| Fairness-aware | Service fairness | Increase priorities for phases not recently selected and reduce preference for a phase retained for a long period. |