Vision--Language Signals in Constrained RL: Safety Gains Without Anticipation
Organizations: Iowa State University Ames, Iowa, USA
Abstract
Safe reinforcement learning seeks policies that maximise task performance while satisfying safety constraints. In driving benchmarks, however, collision costs typically appear only at the time of collision, providing no advance warning of an approaching hazard. Frozen vision--language models can provide dense semantic feedback, yet it remains unclear whether their scores anticipate collisions and which component drives an observed safety improvement. Episodic cost can also favour policies that make little task progress. To address these gaps, we propose VLM-Safe-RL, a framework that integrates frozen CLIP signals into PPO-Lagrangian through reward shaping and an augmented multiplier update. On MetaDrive Hard, which combines the densest traffic with the largest map, the catastrophe rate falls from 31.6% to 19.4%. FormulaOne-L2 analysis finds no evidence that the CLIP signals anticipate collisions and shows that the VLM term has a negligible effect on the Lagrange multiplier. These findings show a conditional reduction in observed catastrophe rate without evidence of collision anticipation.
Figures & tables
| Environment | Cat. % | Viol. % | Cat. 95% CI | ||
|---|---|---|---|---|---|
| PPOLag | VLM Conf | PPOLag | VLM Conf | (pp) | |
| MetaDrive Easy | 14.0 | 32.7 | 18.0 | 38.7 | |
| MetaDrive Medium | 26.0 | 28.0 | 32.8 | 36.0 | |
| MetaDrive Hard | 31.6 | 19.4 | 39.2 | 26.0 | |
| Bullet Car-Reach | 13.2 | 12.7 | 20.8 | 19.8 | |
| F1-L2 † | 2.5 | 8.0 | 17.5 | 18.0 | n/a |
Appendix figures & tables33 assets
Supplementary material from the paper’s appendix.
Appendix
| Group | Hyperparameter | Value |
|---|---|---|
| Optimisation | Total timesteps | (FormulaOne, MetaDrive, Bullet), (Bullet long-horizon runs) |
| Steps per epoch | ( epochs for -step runs, epochs for the Bullet -step runs) | |
| Discount | ||
| GAE | ||
| Learning rate (actor / critic) | / , CPO-family runs have no actor learning rate and a critic learning rate | |
| Linear learning-rate decay | on |
| Env. | Method | Runs | Time (h) | Rollout (h) | Steps/s | vs. ref. | |
| FormulaOne-L2 | PPOLag, 10 iter. | none | 3 | 18.6 [18.4, 18.7] | 18.4 | 14.9 | |
| PPOLag, 40 iter. | none | 2 | 19.2 [19.2, 19.2] | 18.2 | 14.5 | ref. | |
| CPPOPID | none | 3 | 19.2 [19.1, 19.2] | 18.2 | 14.5 | ||
| PPOLag-Dec. | 1 | 2 | 23.1 [22.9, 23.2] | 22.2 | 12.0 | ||
| VLMPPOLag | 1 | 3 | 23.4 [23.0, 24.0] | 22.5 | 11.9 | ||
| Conf, prior | 1 | 3 | 23.3 [23.1, 23.5] | 22.3 | 11.9 |
| Positive prompts ( ) | Negative prompts ( ) |
|---|---|
| FormulaOne | |
| “the racecar is centered on the track and driving safely” | “the racecar is about to crash into the barrier” |
| “the racecar is following the track perfectly with smooth steering” | “the racecar is off the track and unsafe” |
| “the racecar is making progress along the racing line” | “the car is colliding with obstacles” |
| “the car is driving efficiently without collisions” | “the car is driving in the wrong direction” |
| MetaDrive | |
| SafeVLA ( Zhang et al., 2025 ) | VLM-RL ( Huang et al., 2024 ) | CPO-CLG † ( Huang et al., 2024 ) | Ours | |
|---|---|---|---|---|
| Safety formulation | CMDP | none | CMDP | CMDP |
| CLIP scoring | N/A | CLG difference | CLG difference | separate cosines |
| VLM role | policy | reward | reward | reward and dual mean |
| Lagrange multiplier | yes | no | no | yes |
| VLM term in multiplier | no | no | no | yes (negligible) |
| Confidence gating | no | no | no | yes |
| Logged | Bonus | Environment return | |||
|---|---|---|---|---|---|
| Arm | final | final | last-5 | last-10 | final |
| CPO-Coupled | 21.62 | 21.28 | 0.381 | 0.367 | 0.337 |
| CPO-Decoupled | 63.88 | 63.63 | 0.398 | 0.318 | 0.242 |
| PPOLag-Decoupled | 63.77 | 63.54 | 0.318 | 0.355 | 0.228 |
| VLMPPOLag | 63.83 | 63.50 | 0.319 | 0.397 | 0.328 |
| PPOLag (unmatched) | 0.74 | 0 | 0.537 | 0.509 | 0.744 |
| Group 1 | Group 2 | ||||||
| Environment cost | |||||||
| VLMPPOLag Conf ∗ | VLMPPOLag | ||||||
| VLMPPOLag Conf ∗ | PPOLag-Decoupled | ||||||
| VLMPPOLag Conf ∗ | PPOLag-RND | ||||||
| VLMPPOLag | PPOLag-Decoupled | ||||||
| Logged return (augmented for VLM-shaped arms) | |||||||
| Method | (augmented) | final | last-10 | Over budget | |
|---|---|---|---|---|---|
| PPOLag-Decoupled | 15 | 63.8 | 33.8 | 42.5 | 1/1 |
| PPOLag-Decoupled | 25 | 63.8 | 40.7 | 43.8 | 3/3 |
| PPOLag-Decoupled | 35 | 63.9 | 49.6 | 46.5 | 1/1 |
| VLMPPOLag Conf ∗ | 25 | 31.8 | 22.4 | 22.5 | 1/5 |
| Arm | Seed | Path | Progress | Cost | Return | Viol. % | Cat. % |
|---|---|---|---|---|---|---|---|
| PPO (no VLM) | 42 | 5.278 | 2.265 | 346.30 | 1.522 | 74 | 68 |
| 123 | 6.902 | 2.611 | 357.48 | 1.535 | 82 | 78 | |
| 456 | 4.889 | 2.006 | 322.30 | 1.292 | 66 | 58 | |
| CPO (no VLM) | 42 | 2.977 | 0.673 | 41.44 | 0.097 | 18 | 12 |
| 123 | 2.317 | 0.498 | 50.58 | 0.053 | 12 | 10 | |
| 456 | 1.393 | 0.430 | 24.68 | 0.177 | 12 | 8 |
| Family A | Family B | |||
| Seed | Final | Last-10 | Final | Last-10 |
| 42 | 17.75 | 16.14 | 21.37 | 21.57 |
| 123 | 36.37 | 22.39 | 28.16 | 21.58 |
| 456 | 21.70 | 32.63 | 13.61 | 12.13 |
| 789 | 23.07 | 24.11 | 27.00 | 21.45 |
| 1024 | 13.26 | 17.43 | 10.51 | 12.65 |
| L0 | L1 | L2 | ||||
|---|---|---|---|---|---|---|
| Method | ||||||
| FOCOPS | 1.13 0.39 | 0.0 0.0 | 0.26 0.08 | 45.1 17.7 | 0.35 0.33 | 27.6 10.2 |
| CUP | 1.76 0.36 | 0.0 0.0 | 0.38 0.17 | 46.9 10.1 | 0.33 0.09 | 33.0 5.8 |
| P3O | 1.69 0.59 | 0.0 0.0 | 0.04 0.35 | 133.0 29.9 | 0.00 0.18 | 58.7 33.0 |
| Method | env | env | Viol. % | Cat. % | Stages | Path | Progress | |
|---|---|---|---|---|---|---|---|---|
| PPO | 3 | 1.450 0.111 | 342.03 14.68 | 74.0 | 68.0 | 0/7 | 5.69 0.87 | 2.294 0.248 |
| CPO | 3 | 0.109 0.051 | 38.90 10.73 | 14.0 | 10.0 | 0/7 | 2.23 0.65 | 0.534 0.102 |
| PPOLag (unmatched) | 3 | 0.626 0.093 | 85.77 15.72 | 22.7 | 19.3 | 0/7 | 3.47 1.31 | 1.045 0.196 |
| PPOLag-Decoupled † | 2 | 0.231 0.040 | 21.64 5.82 | 14.0 | 6.0 | 0/7 | 1.91 0.10 | 0.566 0.022 |
| VLMPPOLag | 3 | 0.240 0.242 | 22.56 7.56 | 16.0 | 7.3 | 0/7 | 2.14 0.33 | 0.547 0.171 |
| VLMPPOLag Conf | 5 | 0.125 0.059 | 15.71 5.52 | 13.2 | 4.4 | 0/7 | 2.31 0.41 | 0.395 0.082 |
| Configuration | (augmented) | Viol. | |
|---|---|---|---|
| Prior-symmetric gate | 48.4 | 30.5 | 1/3 |
| Ungated, | 63.8 | 40.2 | 2/3 |
| Ungated, | 63.8 | 40.7 | 3/3 |
| Coupled CPO (unmatched) | 21.6 | 32.4 | 3/3 |
| VLM-free PPOLag (unmatched) | 0.7 | 55.8 | 2/3 |
| F1-L0 (final) | F1-L1 (final) | F1-L2 (last 10) | ||||
| Seed | ||||||
| 42 | 42.1 | 0.0 | 29.6 | 6.7 | 24.6 | 16.1 |
| 123 | 41.2 | 0.0 | 39.2 | 20.4 | 33.6 | 22.4 |
| 456 | 36.6 | 0.0 | 24.2 | 44.4 | 52.5 | 32.6 |
| 789 | 41.5 | 0.0 | 34.8 | 11.2 | 15.8 | 24.1 |
| 1024 | 61.2 | 0.0 | 39.9 | 21.4 | 32.6 | 17.4 |
| Method | Seed | Mean cost | Viol% | Cat% | |
| VLMPPOLag+Conf (family A) | 42 | 15.0 | 10.0 | ||
| 123 | 20.0 | 15.0 | |||
| 456 | 30.0 | 10.0 | |||
| 789 | 10.0 | 5.0 | |||
| 1024 | 15.0 | 0.0 | |||
| mean | 18.0 | 8.0 |
| Cell | median margin | median ( , ) | median (calibrated) |
|---|---|---|---|
| F1-L0 | 0.011 | 0.50 | 0.16 |
| F1-L1 | 0.018 | 0.72 | 0.26 |
| F1-L2 | 0.025 | 0.84 | 0.28 |
| MD-Easy (s42) | 0.046 | 0.98 | 0.08 |
| MD-Easy (s123) | 0.046 | 0.98 | 0.30 |
| MD-Easy (s456) | 0.026 | 0.86 | 0.05 |
| Signal | Raw | Within-episode standardized |
|---|---|---|
| 0.172–0.320 | 0.371–0.497 | |
| 0.199–0.452 | 0.348–0.462 | |
| 0.149–0.398 | 0.322–0.465 | |
| Margin | 0.173–0.320 | 0.356–0.493 |
| Step index (control) | 0.512–0.614 | |
| prior-symmetric | calibrated | (calibrated prior) | ||||
| Seed | ||||||
| 42 | 48.0 | 21.6 | 24.6 | 16.1 | ||
| 123 | 49.8 | 38.2 | 33.6 | 22.4 | ||
| 456 | 46.5 | 28.4 | 52.5 | 32.6 | ||
| Mean std ( ) | 48.1 1.3 | 29.4 6.8 | 36.9 11.7 | 23.7 6.8 | ||
| Within budget | 1/3 | 2/3 | ||||
| Method | Training | Seed | Mean cost | Viol. % | Cat. % |
|---|---|---|---|---|---|
| PPOLag | 1M | 42 | 31.5 | 18.0 | 11.5 |
| PPOLag | 1M | 123 | 39.9 | 19.5 | 14.0 |
| PPOLag | 1M | 456 | 42.5 | 23.0 | 16.0 |
| VLMPPOLag Conf | 1M | 42 | 37.6 | 21.5 | 14.5 |
| VLMPPOLag Conf | 1M | 123 | 32.5 | 21.0 | 12.0 |
| VLMPPOLag Conf | 1M | 456 | 40.9 | 23.5 | 15.5 |
| Level | Method | Seed | Viol. | Cat. | |
|---|---|---|---|---|---|
| Easy | PPOLag | 42 | |||
| Easy | PPOLag | 123 | |||
| Easy | PPOLag | 456 | |||
| Easy | VLM+Conf | 42 | |||
| Easy | VLM+Conf | 123 | |||
| Easy | VLM+Conf | 456 |
| Cat. % | Viol. % | ||||||
|---|---|---|---|---|---|---|---|
| Difficulty | Arm A ( ) | Arm B ( ) | Seed range | Arm A | Arm B | ||
| Easy | 32.7% (3) | 25.2% (5) | pp | 8–50% | 0.54 | 38.7% | 30.4% |
| Medium | 28.0% (5) | 21.2% (5) | pp | 6–44% | 0.43 | 36.0% | 26.8% |
| Hard | 19.4% (10) | 27.6% (5) | pp | 0–48% | 0.29 | 26.0% | 36.8% |
| Method | Level | Seed | ||
| Baselines (no VLM) | ||||
| PPO | L0 | 42 | 2.01 | 0.00 |
| PPO | L0 | 123 | 0.91 | 0.00 |
| PPO | L0 | 456 | 1.95 | 0.00 |
| PPO | L1 | 42 | 1.65 | 273.86 |
| PPO | L1 | 123 | 1.49 | 186.59 |
| Level | Seed | k | k | k | M |
|---|---|---|---|---|---|
| L1 | 42 | 0.0097 | 0.0107 | 0.0132 | 0.0101 |
| L1 | 123 | 0.0103 | 0.0133 | 0.0138 | 0.0127 |
| L1 | 456 | 0.0157 | 0.0199 | 0.0212 | 0.0224 |
| L2 | 42 | 0.0128 | 0.0121 | 0.0126 | 0.0106 |
| L2 | 123 | 0.0140 | 0.0126 | 0.0159 | 0.0176 |
| L2 | 456 | 0.0230 | 0.0247 | 0.0192 | 0.0157 |
| Seed | k | k | k | M |
|---|---|---|---|---|
| 42 | 0.620 | 0.621 | 0.610 | 0.637 |
| 123 | 0.610 | 0.619 | 0.601 | 0.616 |
| 456 | 0.618 | 0.640 | 0.641 | 0.630 |
| Arm 1 | Arm 2 | Metric | mean 1 | mean 2 | Bootstrap CI on | Welch | Perm. ( dir.) | |
|---|---|---|---|---|---|---|---|---|
| Qwen2-VL Conf | CLIP Conf | ( ) | ||||||
| CLIP Conf | PPOLag-RND | ( ) | ||||||
| Qwen2-VL Conf | CLIP Conf | logged | ( ) | |||||
| CLIP Conf | PPOLag-RND | logged | ( ) |
| (default) | (3/3) § | (2/3) | (2/3) ‡ |
|---|---|---|---|
| (2 default) | (3/3) | (2/3) ‡ | (2/3) ‡ |
| † This VLM-free cell reports environment return, which is not comparable with the shaped returns in the other cells. § One run stopped at steps. Excluding it gives across the remaining two seeds. | |||
| Backbone | Params | Viol. | ||
| CLIP ViT-B/32 (calibrated, family A) | 151 M | {\color[rgb]{0,0,1}22.5}\pm 5.9 | 1/5 | |
| CLIP ViT-L/14 | 428 M | 1/3 | ||
| One ViT-B/32 scoring call takes ms on an A100. ViT-L/14 latency was not benchmarked. | ||||