An estimated 1 billion people worldwide live with vision impairment, yet current vision-language models (VLMs) produce descriptions too vague for safe navigation by blind and low-vision (BLV) users. Large VLMs can generate high-quality audio-description-compliant narrations but cannot run on mobile devices; small VLMs offer competitive latency but lack spatial detail, directional cues, and hazard awareness for navigational assistance. We present Smol-VL-BLV, a compact VLM for blind and low-vision users that closes this gap using a 500M decoder transformer model and two post-training mechanisms: (1) teacher-student distillation and (2) Group Relative Policy Optimization (GRPO) with a composite BLV reward targeting directional language, metric distances, and hazard detection. Because multi-stage post-training can induce catastrophic forgetting, we add a lightweight finetuning stage after the last stage GRPO finetuning to recover general descriptive quality while preserving BLV-specific spatial grounding. Our best model substantially outperforms the baseline across various benchmarks, including tasks: VQA, BLV captioning, OCR, and latency. Compared with the baseline for relative improvement, it improves the Spatial score gain of 19.3%, and the Social score gain of 14.8%. It also increases OCR-Bench by 101.5%, and raises TextVQA accuracy by 44.2%. These results show that BLV-focused post-training improves both accessibility-specific spatial grounding and general visual-text reasoning. Deployed on a mid-range Android smartphone via Mixed-Precision Quantization, the model remains approx. 450 MB and runs entirely on-device, offline and without network dependency, generating descriptions with latency dependent on host hardware capabilities. Our model, dataset, and code is publicly released at https://smol-vl-blv.github.io/Smol-VL-BLV-website/
Figures & tables
Figure 2: Overview of the post-training pipeline. LUV keyframes and a BLV-oriented text prompt are encoded separately through the vision and text encoders, aligned through lightweight adapter layers, and merged in a fusion block via token concatenation. The decoder, fine-tuned with LoRA while keeping selected backbone components frozen, generates spatially grounded BLV descriptions that include distances, directions, objects, obstacles, etc.
Figure 3: Coverage of BLV-relevant attributes in teacher-generated captions. The captions consistently include core navigation cues such as directions, distances, and spatial layout, while also covering social context, ambience, obstacles, actions, and moving hazards. Lower coverage for step changes indicates an important remaining challenge for BLV-oriented scene description.
Component
Criterion
Score
Caution
Both ref. & gen. contain it
+0.40
Ref. has, gen. missing
−0.30
Gen. has, ref. missing
−0.10
Directional
{left, right, center, ahead, behind}
+0.08 /word
(max +0.24 )
Distances
≥ 1 metric distance match
+0.20
Table 1: GRPO reward components. Range: [−0.50,+1.06] . † “appears to be,” “I cannot,” etc. ‡ 14 scene-type terms.
Model
Params
OCR ANLS
VQA Accuracy
VQA ANLS
SmolVLM2-256M base
256M
16.49%
1.08%
3.09%
SmolVLM2-500M base
500M
32.54%
37.75%
47.52%
SmolVLM2-2.2B base
2.2B
25.44%
27.51%
29.64%
LLaVA-1.5-7B
7B
28.78%
47.79%
62.96%
PaliGemma-3B
3B
69.76%
76.09%
87.36%
Qwen2-VL-2B
2B
86.10%
80.73%
89.24%
Table 2: Comparison of OCR and VQA performance across different vision-language models. Our 500M model substantially improves over the SmolVLM2-500M baseline and remains competitive with larger models.
Model
Params
Spatial
Social
OCR
VQA
GFLOPs
(MCF sub-dims)
PaliGemma2-3B
3B
1.15
1.02
7.8
14.3
10,543
moondream2
1.8B
1.00
1.00
–
–
–
SmolVLM2-2.2B
2.2B
3.49
3.45
25.4
27.5
125,406
SmolVLM2-500M(base)
0.5B
3.21
3.30
32.5
37.8
34,855
Smol-VL-BLV (Ours)
0.5B
3.83
3.79
65.5
54.5
34,855
Table 3: Overall comparative analysis ( n=47 , GPU server). Spatial and Social are MCF sub-dimension scores (1–10 scale). OCR = OCR-Bench ANLS (%). VQA = TextVQA Accuracy (%). ms = mean latency.
Metric
Paper Baseline
SFT
GRPO
GRPO + sft patch
Quantization
Vivo Y27, INT8
A55, Q4_K_M
A55, Q4_K_M
A55, IQ4_NL
Model (LM)
–
290 MB
290 MB
251 MB
Total on-device
–
481 MB
481 MB
442 MB
RAM load time
–
504 ms
488 ms
291 ms
TTFT (prompt eval)
–
∼35 s
35.3 s avg
25.1 s avg
Generation speed
13.55 tok/s
17.1 tok/s
19.1 tok/s
39.3 tok/s
Table 4: On-device performance comparison across model versions and hardware configurations
Metric
Cloud (T4)
A55 (CPU)
Mac (Metal)
TTFT / Prefill
1.39 s
25.1 s
1.95 s
Generation Time
0.31 s
1.4 s
0.46 s
Total Latency
1.70 s
27.1 s
2.41 s
New Tokens
58
60
61
Gen. Speed (tok/s)
187.42
42.9
132.6
Latency/Token (ms)
5.34
23.3
7.54
Table 5: Cross-hardware latency for the same quantized model. The A55’s disproportionate latency reflects CPU-only fallback from missing NPU/GPU driver support (Appendix A.3 ), not a property of the model itself.
Metric
A (Base)
B (SFT)
C (+GRPO)
D (+Patch)
Δ
BLEU-1
19.02
18.08
22.55
45.57
+26.55
BLEU-4
0.74
1.66
1.97
13.53
+12.79
ROUGE-L
10.74
17.66
17.93
30.49
+19.75
METEOR
13.76
15.65
17.68
35.99
+22.23
CIDEr
0.0002
0.0021
0.0018
0.0112
+0.011
Table 6: NLP metrics on the 458-sample final evaluation set. Condition D achieves large gains across all metrics.
Metric
A (Base)
B (SFT)
C (+GRPO)
D (+Patch)
BLV Coverage (%)
Spatial
79
81
97
100
Social
94
94
98
100
Action
68
79
88
62
Ambience
76
91
91
98
BLV Mean
79
86
93
90
Table 7: BLV keyword coverage (458 samples). Condition D achieves 100% coverage on spatial orientation, directions, and distances.
Dimension
A (Base)
B (SFT)
C (+GRPO)
D (+Patch)
LLM Judge (1–10 scale)
MCF Score
1.88
2.26
2.39
2.94
NAF Score
1.84
2.44
2.35
3.57
MCF Sub-Dimensions
Spatial orient.
1.21
1.71
1.85
3.23
Social interact.
1.47
1.62
1.82
2.60
Table 8: LLM judge scores and sub-dimension breakdown (1–10, 458 samples). Condition D achieves the highest scores on 7 of 8 dimensions, with GRPO (C) leading only on ambience.
Method
MCF
NAF
Overall
Base (zero-shot)
3.90
3.98
3.92
SFT v2
4.38
3.98
3.97
DPO
3.89
3.98
3.92
RLAIF-V DPO
3.88
3.96
3.91
GRPO
4.18
4.20
4.19
Table 9: Preference alignment comparison on the balanced evaluation set. Base, SFT v2, DPO, and RLAIF-V DPO are evaluated on 469 samples using a 1–5 judge scale. The DPO and RLAIF-V DPO fall below SFT v2 on MCF.
Figure 4: Qualitative comparison of scene descriptions generated by our model, SmolVLM-500M (Base), and Qwen-3B across four indoor scenes. Our model produces concise, spatially grounded, navigation-focused descriptions, while base models generate generic, observation-style captions without actionable spatial information.
Appendix figures & tables4 assets
Supplementary material from the paper’s appendix.
Appendix
Component
Architecture Detail
Params
Role
Vision Encoder
SigLIP, ViT-style, Patch/14, 384 × 384 input
∼ 86.4M
Pixels → patch embeddings
Vision Projector
Linear projection + LayerNorm
∼ 11.8M
Vision space → LM tokens
Language Model
LLaMA arch, 32 blocks, ctx 8192, vocab 49280
∼ 409.3M
Text generation
Full model
blv_final/ on server
∼ 507.5M
Single artifact
Appendix
Table 10: Neural network component breakdown of the deployed model.
Method
What We Did
Outcome
Q4_K_M
Standard 4-bit K-quant; Apr 25 & May 9 benchmarks.
∼ 340 MB. RSS > 1050 MB with mmproj; quality gain marginal.
Q8_0
Closest-to-F16 quality; evaluated theoretically.
783 MB; rejected—total RAM exceeds 1 GB plus mmproj plus OS.
IQ4_XS
Importance-weighted 4-bit extra-small.
∼ 230 MB; noticeably worse coherence on BLV spatial descriptions.
IQ4_NL (selected)
imatrix from BLV captions; Q5_K on attn layers.
251 MB; TTFT 25.1 s; 39.3 tok/s. Final deployed model.
Q3 & below
IQ3, Q3, Q2 variants; reviewed llama.cpp docs.
Rejected. Sub-4-bit renders small VLMs ( < 1B) incoherent (SPEED-Q).
Appendix
Table 11: Quantization methods explored and outcomes.
Metric
Min
Avg
Max
Model load into RAM
272 ms
291 ms
321 ms
Prompt eval (TTFT)
24.1 s
25.1 s
28.2 s
Generation
1.0 s
1.4 s
1.7 s
Total per frame
26.0 s
27.1 s
30.5 s
Prompt speed (tok/s)
32.7
37.0
38.2
Generation speed (tok/s)
34.2
39.3
43.2
Appendix
Table 12: Per-frame benchmarks for the final IQ4_NL deployment. 2 of 21 frames showed thermal throttling; 19 frames sustained normal speed. No crashes or OOM events.
Multimodal AI, powered by Large Language Models (LLMs) and Vision-Language Models (VLMs), is transforming assistive technologies by enabling simultaneous processing of visual and textual data. This advancement holds significant promise for over 43 million visually impaired and neuro-divergent individuals worldwide who face persistent challenges in navigating indoor and outdoor environments due to limited spatial awareness and insufficient environmental cues. Existing navigation aids often lack comprehensive 3D scene understanding, relying on constrained route-based strategies that hinder user autonomy. In this paper, we introduce a novel end-to-end framework that integrates LLMs, VLMs and digital twin technologies to deliver a spatially cognitive navigation support for visually impaired and neuro-divergent users. Our system captures video input via standard mobile phone cameras, and employs SLAM3R to generate dense 3D point clouds from monocular RGB sequences in real-time. Our custom post-processing algorithm ensures accurate point cloud alignment across multiple viewpoints without requiring predefined reference points. This enhances the capabilities of SpatialLM to produce structured 3D representations, including architectural elements and oriented object bounding boxes. The enriched spatial data is then processed by a locally deployed LLM, which interprets 3D contexts to generate detailed scene descriptions and precise distance measurements between users and surrounding objects. We evaluated our approach across diverse video scenarios featuring various perspectives, looped walking views and captured in multiple environments. The evaluation results demonstrate consistent accuracy in 3D scene interpretation and object localisation, underscoring the potential of our system as a transformative assistive navigation solution that combines advanced visual perception with spatial reasoning
H. Riaz, J. B. Fernandez, I. Mills +3
School of Electronic Engineering, Dublin City University, Dublin, Ireland · Insight Research Ireland Centre for Data Analytics, Dublin, Ireland · South East Technological University, Waterford, Ireland
Visual impairment affects hundreds of millions of people worldwide, severely limiting their ability to navigate urban environments safely and independently. While wearable assistive devices offer a promising platform for real-time hazard detection, existing approaches rely on task-specific vision pipelines that lack flexibility and generalizability. In this work, we propose an event map framework based on visual question answering that leverages Vision-Language Models (VLMs) for pedestrian scene description and hazard identification across diverse real-world environments, using a three-level hierarchical query structure to enable fine-grained scene understanding without task-specific retraining. Model responses are aggregated into a weighted risk scoring system that maps street segments into four discrete safety categories, producing navigable risk-aware event maps for route planning. To support evaluation and future research, we introduce a geographically diverse dataset spanning 20 cities across six continents, comprising over 800 annotated images and 18,000 answered questions. We benchmark four VQA architectures -ViLT, LLaVA, InstructBLIP, and Qwen-VL- and find that generative Multimodal Large Language Models (MLLMs) substantially outperform classification-based approaches, with Qwen-VL achieving the best overall balance of precision and recall. These results demonstrate the viability of MLLMs as a flexible and generalizable foundation for assistive navigation systems for visually impaired people.
Antoni Valls, Jordi Sanchez-Riera
Institut de Rob`otica i Inform`atica Industrial, CSIC-UPC, Llorens i Artigas 4-6, 08028 Barcelona, Spain
Vision-Language Models (VLMs) perform well on commonsense reasoning tasks but struggle with visual spatial reasoning. Most existing solutions introduce extra 3D prior inputs or external spatial encoders, which increase complexity and degrade the underlying VLMs' general-purpose capabilities after spatial fine-tuning. To this end, we propose a parameter-efficient \textit{\textbf{Spatio}-vision \textbf{L}anguage \textbf{M}odels (SpatioLM)}, that enhances spatial intelligence without extra 3D prior inputs or third-party spatial encoders. Concretely, we design a plug-and-play and non-invasive spatio-vision module that elicits the spatial knowledge inherent in VLMs. Furthermore, we innovatively leverage pseudo depth and camera information as supervision to guide the model in learning physically coherent representations. Extensive experiments show that SpatioLM achieves significant improvements in diverse tasks, including spatial perception and understanding while effectively limiting the degradation of general capabilities. Notably, the model achieves an impressive score of 71.6 on the VSI-Bench (the first model to surpass 70). In addition, it attains competitive performance when transferred to embodied manipulation tasks. Code is available at \href{https://github.com/xiaomi-research/spatio-lm}{\faGithub~spatio-lm}.
Jing Wu, Jianhua Wu, Jiayi Guan +5
Xiaomi EV, Beijing, China · College of Automotive and Energy Engineering, Tongji University, Shanghai, China · Independent Researcher