cs.AISep 30, 2026
SaveCompletion-Aware Cross-Fidelity Offline-to-Online Reinforcement Learning for Multi-Line Bus Holding
Organizations: Central South University, Changsha, Hunan, China
Abstract
Exploratory reinforcement learning (RL) on an operating bus fleet is impractical,while policies trained only from historical data cannot acquire new experience. Hybrid Offline-and-Online (H2O) RL combines fixed target replay with simulator interaction, but the inexpensive online simulator can differ from the target in transition and event-duration dynamics. We study this cross-fidelity problem for multi-line bus holding and address a failure mode in which lower generalized passenger time coexists with incomplete passenger journeys.
Figures & tables
| Frozen learning evidence | Adaptive development | Fresh formal confirmation |
| Target-SUMO replay; calibrated online simulator; five IQL-initialized H2O+ actors; 25,000 simulator events; fixed endpoint checkpoints. | Terminal actor taper and deterministic hold transforms were developed over successive preregistered rounds. Development outcomes informed the final completion-safety reserve and therefore are not confirmatory. | The fixed candidate and fixed direct parent were evaluated on 30 untouched SUMO cells: ten independent blocks crossed with three demand levels. Exact paired block tests used all five training seeds within each block. |
Table 1: Separation of learning, adaptive development, and formal confirmation. Only the final column supplies confirmatory evidence.
Figure 1 : Topology of the twelve directional bus services represented in both the target SUMO environment and the calibrated online simulator . The two environments share stop sequences, timetables, capacities, and passenger origin-destination journeys; their traffic and event-propagation mechanisms differ. Line 7X is drawn more heavily only for legibility; policy control covers all 12 services.
| Policy | Gen. passenger s/departed | Completion | Unfinished |
|---|---|---|---|
| Direct parent | 2,484.86 | 0.96185 | 655.42 |
| Reserve-augmented candidate | 2,392.73 | 0.96650 | 578.20 |
| Difference |
Table 2: Formal endpoint means for the preregistered candidate and its direct behavior parent. The difference is candidate minus parent.
| Demand | (s/departed) | (passengers) | |
|---|---|---|---|
| 0.75 | |||
| 1.00 | |||
| 1.25 |
Table 3: Candidate-minus-parent formal differences by demand multiplier. Negative generalized time and unfinished counts are favorable; positive completion is favorable.
| Algorithm label | Gen. passenger s/departed | Completion | Unfinished passengers |
|---|---|---|---|
| Reserve-augmented candidate | 2,392.73 | 0.96650 | 578.20 |
| Direct behavior parent | 2,484.86 | 0.96185 | 655.42 |
| Raw v14 terminal-quarter actor | 2,623.19 | 0.96453 | 610.03 |
| Base H2O+ | 2,696.98 | 0.96288 | 637.41 |
| Online-only SAC | 2,593.85 | 0.96232 | 648.31 |
| Delayed actor (v10) | 2,532.20 | 0.95894 | 702.57 |
Table 4: Descriptive formal means for the fixed candidate, direct parent, raw actor, and lineage controls. Bold denotes the best displayed mean.
Appendix figures & tables3 assets
Supplementary material from the paper’s appendix.
Appendix
| Parameter | Value |
|---|---|
| Canonical state width | 229 |
| Maximum event slots | 16 |
| Slot embedding and fused widths | 128 and 256, respectively |
| Actor/critic/value hidden widths | 256, 256 |
| Critic heads | 2 |
| Maximum hold | 60 s |
Table B.1: Final actor-training and architecture values.
| Diagnostic | Mean | Minimum | Maximum |
|---|---|---|---|
| Training joint AUC | 0.9477 | 0.9370 | 0.9580 |
| Held-out joint AUC | 0.4520 | 0.4366 | 0.4747 |
| Held-out raw ESS fraction | 0.0615 | 0.0100 | 0.1530 |
Table B.2: Domain-ratio diagnostics across five sealed actor-training seeds.
| Seed | (s/departed) | (passengers) | |
|---|---|---|---|
| 6201 | |||
| 6202 | |||
| 6203 | |||
| 6204 | |||
| 6205 |
Table B.3: Formal candidate-minus-parent differences by training seed.