It Is Not Seeing the Hazard: A Frozen Vision-Language Safety Score Measures Its Caption Bank
Organizations: Iowa State University Ames, Iowa, USA
Abstract
Frozen vision-language models increasingly provide safety signals for reinforcement learning. Their use assumes that similarity to language describing danger indicates the hazard itself. Yet policy return and collision rate cannot reveal whether a score detects hazards or responds to correlated features of the scene. VLM-based methods have reported gains in driving and safe-RL benchmarks by converting image-text similarity into rewards, costs, or confidence weights. Such signals promise to reduce reliance on manually designed feedback. They may also reflect prompt structure, embedding geometry, or camera viewpoint, leaving their safety meaning unverified. To address this gap, we present a controlled evaluation of a frozen CLIP prompt-margin safety score. We apply the score to trajectories generated by policies that never receive it, match pre-contact observations to contact-free observations with comparable hazard geometry, and vary the captions, encoder, and camera view. Across three policies, 180 episodes, and 130 isolated contact onsets, the score decreases for about twenty steps before contact. Mechanism controls indicate that the score mainly tracks resemblance to the scene shared by its captions and changes with caption separation and camera view. A constant-confidence control retains the lower catastrophe-rate point estimate, so policy gains do not establish hazard perception.
Figures & tables
| environment | egocentric | chase | overhead near | overhead far |
|---|---|---|---|---|
| FormulaOne (10 policies) | ||||
| CarButton1 (506 onsets) | ||||
| PointGoal1 (156 onsets) |
| scorer | captions | safe group | danger group | margin | redraws |
|---|---|---|---|---|---|
| CLIP ViT-B/32 | 8-caption bank | 27.5/100% | |||
| CLIP ViT-L/14 | 8-caption bank | 100/100% | |||
| CLIP ViT-B/32 | 2 group labels | n/a | |||
| Qwen2-VL-7B | 2 group labels | 51/86.5% | |||
| Qwen2-VL-7B | 8-caption bank | 72/99% |
Appendix figures & tables15 assets
Supplementary material from the paper’s appendix.
Appendix
| scorer | prompt set | what it varies |
|---|---|---|
| CLIP ViT-B/32 | eight-caption bank | the deployed configuration |
| CLIP ViT-L/14 | eight-caption bank | backbone, within the contrastive class |
| Qwen2-VL-7B | two group labels | model class and prompt set together |
| Qwen2-VL-7B | eight-caption bank | model class, prompt set held fixed |
| CLIP ViT-B/32 | two group labels | prompt set, encoder held fixed |
| CLIP ViT-B/32 | three rewrites of the positive group | what the captions claim |
| matching specification | matched | 95% CI | |
|---|---|---|---|
| none (raw) | n/a | n/a | |
| distance | 100% | ||
| distance bearing | 99% | ||
| distance | 99% | ||
| frontal distance | 99% | ||
| distance bearing | 98% |
| deployed positive/negative grouping | |
| rank among all 70 partitions | 3 |
| permutation | |
| partitions with a stronger effect | 2 |
| mean over all 70 partitions | |
| common mode (mean of all eight captions, no contrast) | |
| as a fraction of the deployed margin | 73% |
| stratum | 95% CI | Ep. | Redraws | |
|---|---|---|---|---|
| nearest hazard near ( ) | 38 | 10% | ||
| nearest hazard far ( ) | 44 | 100% | ||
| something within frontal cone | 67 | 100% | ||
| nothing within frontal cone | 8 | 99% |
| arm | seeds | catastrophe | violation | |||
|---|---|---|---|---|---|---|
| baseline | no VLM at all | |||||
| deployed | CLIP, current frame | |||||
| constant | fixed at | |||||
| yoked | real trace, wrong episode |
| view | description | CV | 95% CI | |
|---|---|---|---|---|
| vision | first-person, head-mounted | 0.224 | ||
| track | chase camera following agent | 0.729 | ||
| fixednear | static overhead, near | 0.436 | ||
| fixedfar | static overhead, far | 0.661 |
| encoder | within | within | across | centroids | residual | 2 labels |
|---|---|---|---|---|---|---|
| CLIP ViT-B/32 | 0.883 | 0.898 | 0.883 | 0.962 | 0.274 | 0.846 |
| CLIP ViT-L/14 | 0.808 | 0.845 | 0.794 | 0.913 | 0.408 | 0.817 |
| caption set | example | band effect [95% CI] | corr. |
|---|---|---|---|
| deployed safe | “centered on the track and driving safely” | 1.000 | |
| deployed danger | “about to crash into the barrier” | 0.897 | |
| negated safe | “not centered … and not driving safely” | 0.951 | |
| absurd | “the racecar is made of cheese” | 0.742 | |
| off-topic | “a bowl of fruit on a wooden table” |
| scorer | captions | danger group | margin |
|---|---|---|---|
| CLIP ViT-B/32 | 8-caption bank | 27.5% | 100% |
| CLIP ViT-L/14 | 8-caption bank | 100% | 100% |
| Qwen2-VL-7B | 2 group labels | 51% | 86.5% |
| Qwen2-VL-7B | 8-caption bank | 72% | 99% |