cs.CVSep 29, 2026

Caption-Mediated Perceived-Safety Estimation for Pedestrian Routing

Authors: Simon Parkinson, Paloma Liu, Wei Zheng, Mohammadreza Sheikhfathollahi

Organizations: School of Computing and Mathematics, Manchester Metropolitan University, Manchester M1 5GD, U.K. · University of Huddersfield, Huddersfield HD1 3DH, U.K. · The University of Manchester, Manchester M13 9PL, U.K.

Abstract

This paper presents an explainable approach to pedestrian routing, in which perceived safety is estimated from street-level imagery through an explicit natural-language intermediate representation. A vision--language model caption is generated and stored before any scoring is undertaken, and the perceived-risk class is derived entirely from structured features of that stored text, so that every segment score remains inspectable by the user. Nine captioning conditions across five model families are benchmarked against a direct Contrastive Language--Image Pre-training (CLIP) image-embedding baseline under an identical downstream pipeline, and the caption-mediated representation is found to reach parity with the image embedding rather than to trail it. The approach was deployed over 654,115 images covering 36 electoral wards in two locations in Northern England (Manchester and Huddersfield). Independent field validation against 3,669 locally collected ratings of 494 images across 70 participant sessions established agreement that is statistically significant but modest, at r=0.262r=0.262, against a measured noise ceiling of 0.737 imposed by disagreement between raters. A single-use confirmatory test then found that a pipeline 44% stronger on the supervised benchmark did not produce measurable improvement in the field (r=0.250r=0.250, p=0.84p=0.84), so the benchmark gains did not predict the deployment gains in this case. Routing behaviour varies systematically with journey length. There is negligible change below 1,km, reaching a median increase of 12.78% in low-risk route length for a median detour of 2.73% on journeys of 3 to 6 km.

Figures & tables

Explore similar work

CardsList
  1. PEDESTRIANQA: A Benchmark for Vision-Language Models on Pedestrian Intention and Trajectory Prediction

    May 23, 2026Naman Mishra, Shankar Gangisetty, C. V. JawaharPedestrian

  2. Safety-oriented sidewalk and road segmentation for smartphone-based assistive navigation

    Jul 23, 2026Hakan Calim, Anamaria Dumitrescu, Adarsh Bhandary Panambur +2Pedestrian DetectionEfficient Semantic Segmentation

  3. Can VLMs Reliably Assess Sidewalk Accessibility Attributes from Pedestrian-Level Imagery?

    Sep 15, 2026Seung Jae Lieu, Diego Morra, Chiara Cadoni +3Pedestrian DetectionAccessibility