Caption-Mediated Perceived-Safety Estimation for Pedestrian Routing
Authors: Simon Parkinson, Paloma Liu, Wei Zheng, Mohammadreza Sheikhfathollahi
Organizations: School of Computing and Mathematics, Manchester Metropolitan University, Manchester M1 5GD, U.K. · University of Huddersfield, Huddersfield HD1 3DH, U.K. · The University of Manchester, Manchester M13 9PL, U.K.
This paper presents an explainable approach to pedestrian routing, in which perceived safety is estimated from street-level imagery through an explicit natural-language intermediate representation. A vision--language model caption is generated and stored before any scoring is undertaken, and the perceived-risk class is derived entirely from structured features of that stored text, so that every segment score remains inspectable by the user. Nine captioning conditions across five model families are benchmarked against a direct Contrastive Language--Image Pre-training (CLIP) image-embedding baseline under an identical downstream pipeline, and the caption-mediated representation is found to reach parity with the image embedding rather than to trail it. The approach was deployed over 654,115 images covering 36 electoral wards in two locations in Northern England (Manchester and Huddersfield). Independent field validation against 3,669 locally collected ratings of 494 images across 70 participant sessions established agreement that is statistically significant but modest, at r=0.262, against a measured noise ceiling of 0.737 imposed by disagreement between raters. A single-use confirmatory test then found that a pipeline 44% stronger on the supervised benchmark did not produce measurable improvement in the field (r=0.250, p=0.84), so the benchmark gains did not predict the deployment gains in this case. Routing behaviour varies systematically with journey length. There is negligible change below 1,km, reaching a median increase of 12.78% in low-risk route length for a median detour of 2.73% on journeys of 3 to 6 km.
Figures & tables
Fig. 1: The caption-mediated pipeline. The caption at stage two is a stored artefact from which all subsequent judgements derive, rather than a post-hoc explanation of a judgement made elsewhere.
Fig. 2: Study areas and extraction boundaries. (a) Huddersfield: eight wards sampled uniformly at 20 m. (b) Manchester: 28 wards under a two-tier scheme, with the eight inner-city wards at 20 m (blue) and the twenty outer wards at 100 m (orange). Dots are retained panorama locations, so the plotted extent is the actual imagery footprint rather than a nominal administrative boundary; the visibly denser fill of the blue area is the 20 m regime. Ward polygons are Office for National Statistics 2024 boundaries.
Huddersfield
Manchester
Electoral wards
8
28
Road sampling interval
20 m
20 m / 100 m
wards at 20 m
8
8
wards at 100 m
—
20
Retained locations
82,589
80,967
Distinct panoramas
61,981
80,889
TABLE I: Deployment statistics for the two study areas.
Representation
USID
Place Pulse
1
LDA(5) + sentiment
0.698±0.074
0.700±0.021
2
TF-IDF → SVD(6)
0.830±0.038
0.728±0.022
3
Sentence emb. → PCA(6)
0.742±0.041
0.679±0.018
4
Sentence embedding (384-d)
0.829±0.038
0.708±0.017
5
CLIP image embedding (512-d)
0.796±0.024
0.696±0.030
TABLE II: Representation over identical Qwen2.5-VL captions.
Fig. 3: Closing the gap to a direct image embedding. Bars show USID macro-F1 under 10×5 -fold cross-validation over identical Qwen2.5-VL captions; the dashed line and shaded band give the CLIP image-embedding baseline ( 0.796±0.024 ). The grey bar is the topic-and-sentiment encoding used by the deployed system.
Condition
Words
USID
Place Pulse
Qwen2.5-VL-7B
247
0.829±0.038
0.708±0.017
Qwen2.5-VL-3B
214
0.754±0.019
0.710±0.014
All captioners (mean emb.)
—
0.742±0.031
0.661±0.023
BLIP-base (10 beams)
98
0.730±0.036
0.577±0.012
Florence-2-base
39
0.697±0.043
0.613±0.035
Florence-2-large
38
0.688±0.049
0.607±0.033
TABLE III: Captioning conditions under a fixed sentence-embedding representation. Word counts are the median caption length on USID. The two CLIP rows are image-embedding baselines at 10 and 512 dimensions respectively.
Fig. 4: One worked example per predicted class, showing the four artefacts stored for every scored location.
LDA+sent.
TF-IDF
Sent. emb.
CLIP
LDA topic count, USID
K=3
0.696
0.830
0.829
0.796
K=5 (deployed)
0.698
0.830
0.829
0.796
K=10
0.702
0.830
0.829
0.796
USID class boundaries
2.62/3.48 (deployed)
0.698
0.830
0.829
0.796
TABLE IV: Sensitivity to topic count and class boundaries.
Medium (1–3 km)
Long (3–6 km)
Severity scale
Detour
Δ Low
Detour
Δ Low
0, 1, 2 (deployed)
+ 0.56%
+ 1.75
+ 2.06%
+ 7.68
0, 1, 3 (convex)
+ 1.07%
+ 3.01
+ 3.33%
+ 9.49
0, 1.5, 2 (concave)
+ 0.58%
+ 2.02
+ 2.39%
+ 8.75
0, 0.5, 3 (threshold)
+ 1.02%
+ 3.35
+ 2.98%
+ 7.92
Probability-weighted
+ 0.00%
+ 0.00
+ 0.01%
+ 0.00
TABLE V: Routing outcomes under alternative class-to-cost mappings, Manchester.
City
Journey
Identical
Detour
Δ Low
Δ severity
(%)
(%)
(%)
Huddersfield
< 1 km
13.0
+ 0.00
+ 0.00
+ 0.000
1–3 km
0.8
+ 0.05
+ 0.15
− 0.001
3–6 km
0.0
+ 1.42
+ 4.27
− 0.043
Manchester
< 1 km
16.0
+ 0.00
+ 0.00
+ 0.000
1–3 km
7.0
+ 1.06
+ 7.55
− 0.090
TABLE VI: Routing trade-off by journey length.
Fig. 5: Routing behaviour and field agreement. Left and centre: detour incurred and low-risk length gained by the safety-weighted route, by journey length and city, over 3,000 origin–destination pairs. Right: mean per-image field rating by predicted class, with 95% confidence intervals.
Fig. 6: Maximum attainable model–human correlation as a function of ratings per image, derived from the estimated single-rating reliability of 0.138. The shaded region is attainable in principle but unattained by the present classifier. At the density of this study ( k=7.4 ) the ceiling is 0.737.
Pattern Recognition Lab, Friedrich-Alexander-University Erlangen-Nuremberg, Erlangen, Germany · Department of Telecommunications, National University of Science and Technology “Politehnica” Bucharest, Romania