Caption-Mediated Perceived-Safety Estimation for Pedestrian Routing
Authors: Simon Parkinson, Paloma Liu, Wei Zheng, Mohammadreza Sheikhfathollahi
Organizations: School of Computing and Mathematics, Manchester Metropolitan University, Manchester M1 5GD, U.K. · University of Huddersfield, Huddersfield HD1 3DH, U.K. · The University of Manchester, Manchester M13 9PL, U.K.
This paper presents an explainable approach to pedestrian routing, in which perceived safety is estimated from street-level imagery through an explicit natural-language intermediate representation. A vision--language model caption is generated and stored before any scoring is undertaken, and the perceived-risk class is derived entirely from structured features of that stored text, so that every segment score remains inspectable by the user. Nine captioning conditions across five model families are benchmarked against a direct Contrastive Language--Image Pre-training (CLIP) image-embedding baseline under an identical downstream pipeline, and the caption-mediated representation is found to reach parity with the image embedding rather than to trail it. The approach was deployed over 654,115 images covering 36 electoral wards in two locations in Northern England (Manchester and Huddersfield). Independent field validation against 3,669 locally collected ratings of 494 images across 70 participant sessions established agreement that is statistically significant but modest, at r=0.262, against a measured noise ceiling of 0.737 imposed by disagreement between raters. A single-use confirmatory test then found that a pipeline 44% stronger on the supervised benchmark did not produce measurable improvement in the field (r=0.250, p=0.84), so the benchmark gains did not predict the deployment gains in this case. Routing behaviour varies systematically with journey length. There is negligible change below 1,km, reaching a median increase of 12.78% in low-risk route length for a median detour of 2.73% on journeys of 3 to 6 km.
Figures & tables
Fig. 1: The caption-mediated pipeline. The caption at stage two is a stored artefact from which all subsequent judgements derive, rather than a post-hoc explanation of a judgement made elsewhere.
Fig. 2: Study areas and extraction boundaries. (a) Huddersfield: eight wards sampled uniformly at 20 m. (b) Manchester: 28 wards under a two-tier scheme, with the eight inner-city wards at 20 m (blue) and the twenty outer wards at 100 m (orange). Dots are retained panorama locations, so the plotted extent is the actual imagery footprint rather than a nominal administrative boundary; the visibly denser fill of the blue area is the 20 m regime. Ward polygons are Office for National Statistics 2024 boundaries.
Huddersfield
Manchester
Electoral wards
8
28
Road sampling interval
20 m
20 m / 100 m
wards at 20 m
8
8
wards at 100 m
—
20
Retained locations
82,589
80,967
Distinct panoramas
61,981
80,889
TABLE I: Deployment statistics for the two study areas.
Representation
USID
Place Pulse
1
LDA(5) + sentiment
0.698±0.074
0.700±0.021
2
TF-IDF → SVD(6)
0.830±0.038
0.728±0.022
3
Sentence emb. → PCA(6)
0.742±0.041
0.679±0.018
4
Sentence embedding (384-d)
0.829±0.038
0.708±0.017
5
CLIP image embedding (512-d)
0.796±0.024
0.696±0.030
TABLE II: Representation over identical Qwen2.5-VL captions.
Fig. 3: Closing the gap to a direct image embedding. Bars show USID macro-F1 under 10×5 -fold cross-validation over identical Qwen2.5-VL captions; the dashed line and shaded band give the CLIP image-embedding baseline ( 0.796±0.024 ). The grey bar is the topic-and-sentiment encoding used by the deployed system.
Condition
Words
USID
Place Pulse
Qwen2.5-VL-7B
247
0.829±0.038
0.708±0.017
Qwen2.5-VL-3B
214
0.754±0.019
0.710±0.014
All captioners (mean emb.)
—
0.742±0.031
0.661±0.023
BLIP-base (10 beams)
98
0.730±0.036
0.577±0.012
Florence-2-base
39
0.697±0.043
0.613±0.035
Florence-2-large
38
0.688±0.049
0.607±0.033
TABLE III: Captioning conditions under a fixed sentence-embedding representation. Word counts are the median caption length on USID. The two CLIP rows are image-embedding baselines at 10 and 512 dimensions respectively.
Fig. 4: One worked example per predicted class, showing the four artefacts stored for every scored location.
LDA+sent.
TF-IDF
Sent. emb.
CLIP
LDA topic count, USID
K=3
0.696
0.830
0.829
0.796
K=5 (deployed)
0.698
0.830
0.829
0.796
K=10
0.702
0.830
0.829
0.796
USID class boundaries
2.62/3.48 (deployed)
0.698
0.830
0.829
0.796
TABLE IV: Sensitivity to topic count and class boundaries.
Medium (1–3 km)
Long (3–6 km)
Severity scale
Detour
Δ Low
Detour
Δ Low
0, 1, 2 (deployed)
+ 0.56%
+ 1.75
+ 2.06%
+ 7.68
0, 1, 3 (convex)
+ 1.07%
+ 3.01
+ 3.33%
+ 9.49
0, 1.5, 2 (concave)
+ 0.58%
+ 2.02
+ 2.39%
+ 8.75
0, 0.5, 3 (threshold)
+ 1.02%
+ 3.35
+ 2.98%
+ 7.92
Probability-weighted
+ 0.00%
+ 0.00
+ 0.01%
+ 0.00
TABLE V: Routing outcomes under alternative class-to-cost mappings, Manchester.
City
Journey
Identical
Detour
Δ Low
Δ severity
(%)
(%)
(%)
Huddersfield
< 1 km
13.0
+ 0.00
+ 0.00
+ 0.000
1–3 km
0.8
+ 0.05
+ 0.15
− 0.001
3–6 km
0.0
+ 1.42
+ 4.27
− 0.043
Manchester
< 1 km
16.0
+ 0.00
+ 0.00
+ 0.000
1–3 km
7.0
+ 1.06
+ 7.55
− 0.090
TABLE VI: Routing trade-off by journey length.
Fig. 5: Routing behaviour and field agreement. Left and centre: detour incurred and low-risk length gained by the safety-weighted route, by journey length and city, over 3,000 origin–destination pairs. Right: mean per-image field rating by predicted class, with 95% confidence intervals.
Fig. 6: Maximum attainable model–human correlation as a function of ratings per image, derived from the estimated single-rating reliability of 0.138. The shaded region is attainable in principle but unattained by the present classifier. At the density of this study ( k=7.4 ) the ceiling is 0.737.
Pedestrian intention and trajectory prediction are critical for the safe deployment of autonomous driving systems, directly influencing navigation decisions in complex traffic environments. Recent advances in large vision-language models offer a powerful new paradigm for these tasks by combining high-capacity visual understanding with flexible natural language reasoning. In this work, we introduce PedestrianQA, a large-scale video-based dataset that formulates pedestrian intention and trajectory prediction as question-answering tasks augmented with structured rationales. PedestrianQA expresses richly annotated pedestrian sequences, in natural language, enabling VLMs to learn from visual dynamics, contextual cues, and interactions among traffic agents while generating concise explanations of their predictions without needing specialized architectures tailored for each task. Empirical evaluations across PIE, JAAD, TITAN, and IDD-PeD show that finetuning state-of-the-art VLMs on PedestrianQA significantly improves intention classification, trajectory forecasting accuracy, and the quality of explanatory rationales, demonstrating the strong potential of VLMs as a unified and explainable framework for safety-critical pedestrian behavior modeling.
Independent sidewalk mobility is essential for blind and visually impaired pedestrians (BVIPs), yet smartphone-based assistive navigation requires perception models that distinguish walkable sidewalks from adjacent unsafe regions. This study presents a safety-oriented semantic segmentation framework for future mobile guidance. We introduce SENSATION-DS, a chest-height pedestrian-view dataset with 2,752 image-mask pairs and nine-class navigation-relevant taxonomy. External urban and sidewalk datasets were harmonized to this label space, and five segmentation architectures were evaluated using staged target-domain adaptation with mask-conditioned synthetic images and Segment Anything Model 2 (SAM2) pseudo-labels. Models were assessed using mean Intersection over Union (mIoU), road- and sidewalk-specific metrics, Road-as-Sidewalk Error Rate as a proxy false-safe measure, and Android Open Neural Network Exchange benchmarking. Synthetic augmentation generally improved segmentation accuracy, whereas SAM2 pseudo-labels more consistently reduced Road-as-Sidewalk errors. UPerNet-MobileNetV3 achieved the highest offline mIoU (0.715 +/- 0.006), while DeepLabV3Plus-MobileNetV3 achieved the lowest Road-as-Sidewalk Error Rate (0.079) and highest Android runtime at 512x384 (7.383 FPS). These results show that assistive sidewalk perception should be evaluated jointly by segmentation accuracy, proxy false-safe behavior, and smartphone deployment feasibility, while real-world benefit requires validation with BVIP users. This evaluation supports selecting models that balance accurate perception, conservative error behavior, and practical runtime.
Pattern Recognition Lab, Friedrich-Alexander-University Erlangen-Nuremberg, Erlangen, Germany · Department of Telecommunications, National University of Science and Technology “Politehnica” Bucharest, Romania
An important component of urban accessibility, particularly for wheelchair users and people with reduced mobility, is sidewalk compliance with measurable requirements. We test whether effective width, longitudinal slope, cross slope, and pavement condition can be assessed reliably from pedestrian-level imagery using vision-language models (VLMs). We present the first application of sampling-based conformal prediction (CP) for VLM-based accessibility assessment. We evaluate four VLMs on 514 sidewalk images from Seoul, South Korea, with field-measured ground truth. Conformal calibration attains the nominal 90% coverage for all models and attributes, but the calibrated regions differ in informativeness. Effective width yields the most informative estimates, with a mean interval half-width of about 1.0 m for the best model. Since every model overestimates width, asymmetric calibration shortens the intervals by up to 33% at unchanged coverage. Longitudinal slope is marginally informative, cross-slope intervals are too wide to resolve regulatory thresholds, and pavement-condition sets degenerate to all five grades (A-E) for three of the four models. Uncalibrated intervals from raw sampling dispersion cover only 17-47% of field-measured values at a nominal 90% level. Among the images with the most self-consistent responses, these intervals miss the field-measured value in up to 96% of cases. Response self-consistency is therefore not evidence of accuracy, and sampling dispersion cannot be interpreted as uncertainty until it has been calibrated against field-measured ground truth. No quantitative attribute reaches the precision required for general compliance assessment, but CP identifies from calibration data alone which attributes can support screening of segments far from the thresholds. We release the annotated pedestrian-level images and their corresponding field-measured attribute values.
Seung Jae Lieu, Diego Morra, Chiara Cadoni +3
Senseable City Lab Massachusetts Institute of Technology Cambridge, MA 02139