Small yet Assistive: Spatially-Aware Post-Training for Low Vision
Organizations: Indian Institute of Technology Mandi · Mohamed bin Zayed University of Artificial Intelligence
Abstract
An estimated 1 billion people worldwide live with vision impairment, yet current vision-language models (VLMs) produce descriptions too vague for safe navigation by blind and low-vision (BLV) users. Large VLMs can generate high-quality audio-description-compliant narrations but cannot run on mobile devices; small VLMs offer competitive latency but lack spatial detail, directional cues, and hazard awareness for navigational assistance. We present Smol-VL-BLV, a compact VLM for blind and low-vision users that closes this gap using a 500M decoder transformer model and two post-training mechanisms: (1) teacher-student distillation and (2) Group Relative Policy Optimization (GRPO) with a composite BLV reward targeting directional language, metric distances, and hazard detection. Because multi-stage post-training can induce catastrophic forgetting, we add a lightweight finetuning stage after the last stage GRPO finetuning to recover general descriptive quality while preserving BLV-specific spatial grounding. Our best model substantially outperforms the baseline across various benchmarks, including tasks: VQA, BLV captioning, OCR, and latency. Compared with the baseline for relative improvement, it improves the Spatial score gain of 19.3%, and the Social score gain of 14.8%. It also increases OCR-Bench by 101.5%, and raises TextVQA accuracy by 44.2%. These results show that BLV-focused post-training improves both accessibility-specific spatial grounding and general visual-text reasoning. Deployed on a mid-range Android smartphone via Mixed-Precision Quantization, the model remains approx. 450 MB and runs entirely on-device, offline and without network dependency, generating descriptions with latency dependent on host hardware capabilities. Our model, dataset, and code is publicly released at https://smol-vl-blv.github.io/Smol-VL-BLV-website/
Figures & tables
| Component | Criterion | Score |
| Caution | Both ref. & gen. contain it | |
| Ref. has, gen. missing | ||
| Gen. has, ref. missing | ||
| Directional | {left, right, center, ahead, behind} | /word |
| (max ) | ||
| Distances | 1 metric distance match |
| Model | Params | OCR ANLS | VQA Accuracy | VQA ANLS |
|---|---|---|---|---|
| SmolVLM2-256M base | 256M | 16.49% | 1.08% | 3.09% |
| SmolVLM2-500M base | 500M | 32.54% | 37.75% | 47.52% |
| SmolVLM2-2.2B base | 2.2B | 25.44% | 27.51% | 29.64% |
| LLaVA-1.5-7B | 7B | 28.78% | 47.79% | 62.96% |
| PaliGemma-3B | 3B | 69.76% | 76.09% | 87.36% |
| Qwen2-VL-2B | 2B | 86.10% | 80.73% | 89.24% |
| Model | Params | Spatial | Social | OCR | VQA | GFLOPs |
|---|---|---|---|---|---|---|
| (MCF sub-dims) | ||||||
| PaliGemma2-3B | 3B | 1.15 | 1.02 | 7.8 | 14.3 | 10,543 |
| moondream2 | 1.8B | 1.00 | 1.00 | – | – | – |
| SmolVLM2-2.2B | 2.2B | 3.49 | 3.45 | 25.4 | 27.5 | 125,406 |
| SmolVLM2-500M(base) | 0.5B | 3.21 | 3.30 | 32.5 | 37.8 | 34,855 |
| Smol-VL-BLV (Ours) | 0.5B | 3.83 | 3.79 | 65.5 | 54.5 | 34,855 |
| Metric | Paper Baseline | SFT | GRPO | GRPO + sft patch |
|---|---|---|---|---|
| Quantization | Vivo Y27, INT8 | A55, Q4_K_M | A55, Q4_K_M | A55, IQ4_NL |
| Model (LM) | – | 290 MB | 290 MB | 251 MB |
| Total on-device | – | 481 MB | 481 MB | 442 MB |
| RAM load time | – | 504 ms | 488 ms | 291 ms |
| TTFT (prompt eval) | – | s | 35.3 s avg | 25.1 s avg |
| Generation speed | 13.55 tok/s | 17.1 tok/s | 19.1 tok/s | 39.3 tok/s |
| Metric | Cloud (T4) | A55 (CPU) | Mac (Metal) |
|---|---|---|---|
| TTFT / Prefill | 1.39 s | 25.1 s | 1.95 s |
| Generation Time | 0.31 s | 1.4 s | 0.46 s |
| Total Latency | 1.70 s | 27.1 s | 2.41 s |
| New Tokens | 58 | 60 | 61 |
| Gen. Speed (tok/s) | 187.42 | 42.9 | 132.6 |
| Latency/Token (ms) | 5.34 | 23.3 | 7.54 |
| Metric | A (Base) | B (SFT) | C (+GRPO) | D (+Patch) | |
|---|---|---|---|---|---|
| BLEU-1 | 19.02 | 18.08 | 22.55 | 45.57 | +26.55 |
| BLEU-4 | 0.74 | 1.66 | 1.97 | 13.53 | +12.79 |
| ROUGE-L | 10.74 | 17.66 | 17.93 | 30.49 | +19.75 |
| METEOR | 13.76 | 15.65 | 17.68 | 35.99 | +22.23 |
| CIDEr | 0.0002 | 0.0021 | 0.0018 | 0.0112 | +0.011 |
| Metric | A (Base) | B (SFT) | C (+GRPO) | D (+Patch) |
|---|---|---|---|---|
| BLV Coverage (%) | ||||
| Spatial | 79 | 81 | 97 | 100 |
| Social | 94 | 94 | 98 | 100 |
| Action | 68 | 79 | 88 | 62 |
| Ambience | 76 | 91 | 91 | 98 |
| BLV Mean | 79 | 86 | 93 | 90 |
| Dimension | A (Base) | B (SFT) | C (+GRPO) | D (+Patch) |
|---|---|---|---|---|
| LLM Judge (1–10 scale) | ||||
| MCF Score | 1.88 | 2.26 | 2.39 | 2.94 |
| NAF Score | 1.84 | 2.44 | 2.35 | 3.57 |
| MCF Sub-Dimensions | ||||
| Spatial orient. | 1.21 | 1.71 | 1.85 | 3.23 |
| Social interact. | 1.47 | 1.62 | 1.82 | 2.60 |
| Method | MCF | NAF | Overall |
|---|---|---|---|
| Base (zero-shot) | 3.90 | 3.98 | 3.92 |
| SFT v2 | 4.38 | 3.98 | 3.97 |
| DPO | 3.89 | 3.98 | 3.92 |
| RLAIF-V DPO | 3.88 | 3.96 | 3.91 |
| GRPO | 4.18 | 4.20 | 4.19 |
Appendix figures & tables4 assets
Supplementary material from the paper’s appendix.
Appendix
| Component | Architecture Detail | Params | Role |
|---|---|---|---|
| Vision Encoder | SigLIP, ViT-style, Patch/14, 384 384 input | 86.4M | Pixels patch embeddings |
| Vision Projector | Linear projection + LayerNorm | 11.8M | Vision space LM tokens |
| Language Model | LLaMA arch, 32 blocks, ctx 8192, vocab 49280 | 409.3M | Text generation |
| Full model | blv_final/ on server | 507.5M | Single artifact |
| Method | What We Did | Outcome |
|---|---|---|
| Q4_K_M | Standard 4-bit K-quant; Apr 25 & May 9 benchmarks. | 290 MB; TTFT 35 s; 17–19 tok/s. Baseline for comparisons. |
| Q5_K_M | 5-bit K-quant; tested spatial-language coherence gain. | 340 MB. RSS 1050 MB with mmproj; quality gain marginal. |
| Q8_0 | Closest-to-F16 quality; evaluated theoretically. | 783 MB; rejected—total RAM exceeds 1 GB plus mmproj plus OS. |
| IQ4_XS | Importance-weighted 4-bit extra-small. | 230 MB; noticeably worse coherence on BLV spatial descriptions. |
| IQ4_NL (selected) | imatrix from BLV captions; Q5_K on attn layers. | 251 MB; TTFT 25.1 s; 39.3 tok/s. Final deployed model. |
| Q3 & below | IQ3, Q3, Q2 variants; reviewed llama.cpp docs. | Rejected. Sub-4-bit renders small VLMs ( 1B) incoherent (SPEED-Q). |
| Metric | Min | Avg | Max |
|---|---|---|---|
| Model load into RAM | 272 ms | 291 ms | 321 ms |
| Prompt eval (TTFT) | 24.1 s | 25.1 s | 28.2 s |
| Generation | 1.0 s | 1.4 s | 1.7 s |
| Total per frame | 26.0 s | 27.1 s | 30.5 s |
| Prompt speed (tok/s) | 32.7 | 37.0 | 38.2 |
| Generation speed (tok/s) | 34.2 | 39.3 | 43.2 |
| Metric | Q4_K_M (May 9) | IQ4_NL (May 24) | |
|---|---|---|---|
| Total per frame | 38.7 s | 27.1 s | 30% |
| TTFT | 35.3 s | 25.1 s | 29% |
| Generation speed | 19.1 tok/s | 39.3 tok/s | 106% |
| LM file size | 290 MB | 251 MB | 13% |
| Total on-device | 481 MB | 442 MB | 8% |
| Peak RAM | 985 MB | 780 MB | 21% |