Speed in the Blind Spot: An Interpretability Analysis of Dynamic Perception in VLMs for Autonomous Driving
Organizations: Munich University of Applied Sciences, Intelligent Vehicles Lab (IVL), 80335 Munich, Germany
Abstract
Vision-Language Models are increasingly used in autonomous-driving systems, yet their ability to recover dynamic physical state from visual input remains insufficiently characterized. We study velocity understanding as a controlled diagnostic across three tasks: surrounding-agent speed, current ego speed, and short-horizon future ego-speed proposal. On nuScenes, we evaluate open-weight general-purpose and PhysicalAI VLMs, together with the driving-oriented Alpamayo-1.5 Vision-Language-Action model, using multiple input and output formulations. We combine verbal evaluation with temporal perturbations, counterfactual ego-speed hints and linear probes of hidden representations. The tasks exhibit distinct failure modes. Surrounding-agent speed is weakly encoded in an agent-specific form, whereas current ego speed is often internally accessible but poorly verbalized: continuous probes achieve 4.7-5.8 km/h MAE compared with 10.2-16.8 km/h MAE for verbal outputs. Multiple frames provide inconsistent verbal gains to single frame inputs, and frame order is rarely exploited. Under non-optimized QLoRA, task-specific adaptation improves both task-relevant latent speed representations and verbal readout, but continuous surrounding-agent speed estimation remains weak, while most future-speed gains survive frame shuffling, indicating limited temporal grounding. Driving specialized Alpamayo-1.5 shows stronger latent representations for surrounding-agent and future ego speed, while current ego-speed decodability is comparable and substantial probe-verbal gaps remain. Thus, driving specialization can strengthen motion representations but does not guarantee stronger encoding across both scene and ego states or reliable readout. The results show that plausible planning outputs do not necessarily imply reliable recovery or temporal grounding of the underlying dynamic state.
Figures & tables
| Model | Vision encoder | LLM backbone |
|---|---|---|
| Cosmos-Reason2-8B [ 36 ] | SigLIP-2 [ 37 ] | Qwen3-8B |
| Cosmos-Reason1-7B [ 38 ] | Qwen2.5-VL ViT | Qwen2.5-7B |
| Qwen2.5-VL-7B [ 20 ] | Qwen2.5-VL ViT | Qwen2.5-7B |
| Qwen3-VL-8B [ 21 ] | SigLIP-2 [ 37 ] | Qwen3-8B |
| InternVL2-8B [ 23 ] | InternViT-300M-448px | internlm2_5-7b-chat |
| InternVL2.5-8B [ 22 ] | InternViT-300M-448px-V2.5 | internlm2_5-7b-chat |
| T1 surrounding vehicle | T2 ego current | T3 ego future | ||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Model | 1F | 5F | 1F+S | 5F+S | 1F | 5F | 1F+S | 5F+S | 1F | 5F | 1F+S | 5F+S |
| Cosmos-R2 | 22.1 | 21.8 | 27.3 | 29.1 | 22.9 | 36.2 | 65.2 | 63.4 | 46.1 | 44.5 | 66.1 | 61.0 |
| Cosmos-R1 | 43.3 | 42.2 | 44.7 | 44.7 | 31.4 | 33.1 | 57.2 | 62.0 | 45.6 | 55.1 | 64.8 | 72.7 |
| Qwen2.5-VL | 33.5 | 39.1 | 41.6 | 43.2 | 22.7 | 27.8 | 50.2 | 64.8 | 35.3 | 35.3 | 52.9 | 54.3 |
| Qwen3-VL | 42.0 | 30.5 | 37.8 | 39.1 | 40.4 | 40.9 | 78.8 | 82.0 | 51.8 | 51.4 | 61.1 | 69.4 |
| LLaVA-Video | 21.6 | 14.2 | 18.4 | 20.3 | 23.0 | 25.3 | 37.2 | 47.4 | 14.6 | 14.1 | 15.2 | 14.2 |
| Purpose | Split | Scenes | Samples |
|---|---|---|---|
| Benchmark evaluation | Val. | 150 | T1: 4,429; T2: 5,419; T3: 4,519 |
| Probe training | Train | 547 | 6,000 ( 2,000/class;joint-stratified) |
| QLoRA adaptation | Train | 560 | 20,275 |
| T1 surrounding | T2 ego current | T3 ego future | ||||
|---|---|---|---|---|---|---|
| Model | 1F | 5F | 1F | 5F | 1F | 5F |
| Cosmos-R2 | 60.1 | 64.7 | 25.9 | 76.6 | 46.5 | 74.2 |
| Cosmos-R1 | 35.7 | 35.2 | 16.3 | 15.8 | 32.8 | 22.6 |
| Qwen2.5-VL | 37.4 | 37.4 | 16.6 | 27.6 | 33.3 | 40.6 |
| Qwen3-VL | 62.6 | 46.2 | 56.0 | 67.9 | 59.3 | 67.8 |
| LLaVA-Video | 48.6 | 36.6 | 15.2 | 15.0 | 14.0 | 14.0 |
| T1 surrounding | T2 ego current | T3 ego future | ||||
|---|---|---|---|---|---|---|
| Model | 1F | 5F | 1F | 5F | 1F | 5F |
| Cosmos-R2 | 10.8 | 10.5 | 19.0 | 14.2 | 15.2 | 12.0 |
| Cosmos-R1 | 17.6 | 15.4 | 14.9 | 10.2 | 13.4 | 14.3 |
| Qwen2.5-VL | 11.7 | 17.0 | 18.4 | 12.2 | 12.8 | 14.2 |
| Qwen3-VL | 15.6 | 15.5 | 19.0 | 10.4 | 15.5 | 10.6 |
| LLaVA-Video | 11.6 | 11.9 | 19.1 | 18.5 | 15.1 | 14.5 |
| T1 – lead speed | T2 – ego speed | T3 – future speed | |||||||
|---|---|---|---|---|---|---|---|---|---|
| Model | 1F | 5F | 1F+S | 1F | 5F | 1F+S | 1F | 5F | 1F+S |
| Cosmos-R1 | 35.6 ( -8 ) | 49.7 ( +8 ) | 43.4 ( -1 ) | 59.6 ( +28 ) | 82.3 ( +49 ) | 96.6 ( +39 ) | 57.4 ( +12 ) | 66.2 ( +11 ) | 69.2 ( +4 ) |
| Cosmos-R2 | 36.7 ( +15 ) | 44.4 ( +23 ) | 38.8 ( +11 ) | 62.3 ( +39 ) | 75.3 ( +39 ) | 95.4 ( +30 ) | 59.5 ( +13 ) | 64.6 ( +20 ) | 68.6 ( +3 ) |
| Qwen2.5-VL | 38.9 ( +5 ) | 29.1 ( -10 ) | 41.9 ( +0 ) | 60.5 ( +38 ) | 77.4 ( +50 ) | 96.9 ( +47 ) | 56.5 ( +21 ) | 67.5 ( +32 ) | 73.3 ( +20 ) |
| Qwen3-VL | 25.6 ( -16 ) | 41.1 ( +11 ) | 21.8 ( -16 ) | 61.3 ( +21 ) | 77.3 ( +36 ) | 96.9 ( +18 ) | 59.3 ( +7 ) | 61.8 ( +10 ) | 68.5 ( +7 ) |
| LLaVA-Video | 38.1 ( +17 ) | 42.5 ( +28 ) | 41.9 ( +24 ) | 55.8 ( +33 ) | 74.1 ( +49 ) | 96.9 ( +60 ) | 54.1 ( +40 ) | 64.8 ( +51 ) | 73.1 ( +58 ) |
| Model | Task | Setting | Probe | Verbal |
| Qwen3-VL | T1 | Zero-shot | 41.1 | 30.5 |
| LoRA-full | 69.9 | 71.0 | ||
| T2 | Zero-shot | 77.3 | 40.9 | |
| LoRA-full | 89.0 | 90.1 | ||
| LoRA LLM-only | 87.5 | 89.0 | ||
| LoRA VE-only | 85.9 | 87.6 |
| Zero-shot | LoRA-full | ||||
|---|---|---|---|---|---|
| Task | Readout | Ordered | Shuffled | Ordered | Shuffled |
| T2 | Verbal | 40.9 | 40.1 | 89.5 | 86.4 |
| Probe | 77.3 | 76.5 | 89.0 | 83.6 | |
| T3 | Verbal | 51.4 | 51.3 | 72.4 | 72.2 |
| Probe | 66.5 | 61.5 | 71.0 | 69.2 | |
| Main finding | Evidence | Potential mitigation / related finding |
|---|---|---|
| Weak agent-specific encoding (T1). Surrounding-agent speed is not reliably encoded in an agent-specific representation by the VLM foundation models. | Section V - V-A (Table VI ); Section V - V-D (Fig. 8 ); Appendix A (Tables X and XI ); Section VI (Table VII ) | Agent-centric temporal or relative-motion supervision. Alpamayo-1.5’s stronger zero-shot T1 probe and the substantial probe gain of adapted Qwen3-VL suggest that driving- and task-specific supervision can strengthen agent-speed representations (Section VI , Table VII ), while continuous T1 estimation on our non-optimized training remains weak (Fig. 9 ). |
| Representation–verbalization gap (T2). Current ego speed is strongly linearly accessible in hidden states but only incompletely reflected in verbal outputs. | Section IV - IV-A (Table II ); Section V - V-A (Table VI ); Section V - V-C (Fig. 7 ); Section VI (Table VII ) | Improve readout while explicitly supervising temporal motion cues. Task-specific adaptation largely closes the readout gap and increases temporal-order sensitivity (Section VI , Table VII ; Fig. 9 ; Table VIII ). |
| Planning performance can mask weak perceptual grounding (T3). T3 predictions and representations can achieve nontrivial target agreement despite limited sensitivity to chronological motion evidence, yet improve strongly when accurate current ego speed is supplied. Thus, planning scores alone can understate the importance of reliable dynamic-state perception. | Section IV - IV-A (Table II ); Section IV - IV-B (Fig. 3 ) Section IV - IV-E (Fig. 5 ); | Evaluate planning jointly with its perceptual prerequisites. Improving explicit recovery and verification of the current dynamic state may strengthen planning-related predictions, while target accuracy alone should not be interpreted as evidence of perceptual or temporal grounding. |
| Limited use of chronological history. Ordered and shuffled five-frame inputs perform similarly in most zero-shot settings, and rare gains from additional history often survive shuffling. | Section IV - IV-E (Fig. 5 ) | Explicit temporal supervision. Task-specific adaptation produces limited evidence of increased order sensitivity (Section VI , Table VIII ), leaving temporal-order or motion supervision as an open direction. |
| Prompt-supplied ego state is integrated zero-shot and dominates visual cues. Without task-specific training on ego-state inputs, VLMs already use numerical ego-speed hints to improve T2/T3. Counterfactual hints strongly mislead both tasks, showing that prompt-supplied state is readily incorporated but weakly cross-checked against visual evidence. | Section IV - IV-A (Table II ); Section IV - IV-B (Fig. 3 ) | Conflict-aware multimodal fusion and uncertainty calibration. Inconsistencies between prompt-supplied state and visually inferred state should be checked and biased prompts should be rejected when inconsistent. Counterfactual hints provide a direct robustness test for such failures. |
| Formulation-sensitive and biased verbal outputs. Paired categorical and metric predictions can contradict, and several models exhibit persistent class preferences or output collapse. Thus, different output interfaces do not reliably expose a consistent speed estimate. | Section IV - IV-C (paired cross-formulation analysis; Tables II , IV , and V ; Fig. 4 ); Section VI | Cross-format supervision with shared state representations. Task-specific adaptation improves the supervised format but transfers poorly across formulations. Cross-format training and shared state representations can be directions for further study. |
Appendix figures & tables4 assets
Supplementary material from the paper’s appendix.
Appendix
| Model | 5F current | 5F all | |
|---|---|---|---|
| Cosmos-Reason2 | 21.8 | 22.9 | +1.1 |
| Cosmos-Reason1 | 42.2 | 43.2 | +1.0 |
| Qwen2.5-VL | 39.1 | 38.5 | -0.6 |
| Qwen3-VL | 30.5 | 26.1 | -4.4 |
| LLaVA-Video | 14.2 | 13.6 | +0.6 |
| InternVL2.5 | 41.4 | 41.4 | +0.0 |
| Model | Accuracy | Macro-F1 |
|---|---|---|
| Cosmos-Reason2 | 76.2 | 34.0 |
| Cosmos-Reason1 | 69.6 | 48.0 |
| Qwen2.5-VL | 73.6 | 49.3 |
| Qwen3-VL | 78.5 | 43.8 |
| LLaVA-Video | 75.8 | 36.7 |
| InternVL2.5 | 70.9 | 44.8 |
| T1 surrounding | T2 ego current | T3 ego future | ||||
|---|---|---|---|---|---|---|
| Model | 1F | 5F | 1F | 5F | 1F | 5F |
| Cosmos-R2 | 17.4 | 14.5 | 22.7 | 18.1 | 19.4 | 15.7 |
| Cosmos-R1 | 23.9 | 19.2 | 18.7 | 15.5 | 18.1 | 26.5 |
| Qwen2.5-VL | 18.2 | 21.1 | 22.0 | 15.9 | 16.6 | 17.9 |
| Qwen3-VL | 22.3 | 19.8 | 22.7 | 13.7 | 19.5 | 14.0 |
| LLaVA-Video | 18.7 | 18.6 | 22.7 | 22.3 | 19.2 | 18.7 |
| Condition | Target speed | Ego speed |
|---|---|---|
| 1F | 0.27 | 0.29 |
| 1F+S | 0.32 | 0.77 |
| 5F+S | 0.34 | 0.84 |