More Than Meets the Eye? Uncovering the Reasoning-Planning Disconnect in Training Vision-Language Driving Models
Organizations: S-Lab, Nanyang Technological University, Singapore · College of Computing and Data Science, Nanyang Technological University, Singapore
Abstract
Vision-Language Model (VLM) driving agents promise explainable end-to-end autonomy by first producing natural-language reasoning and then predicting trajectory planning. However, whether planning is causally driven by this reasoning remains a critical but unverified assumption. To investigate this, we build DriveMind, a large-scale driving Visual Question Answering corpus with plan-aligned Chain-of-Thought (CoT), automatically generated from nuPlan. Our data generation process converts sensors and annotations into structured inputs and, crucially, separates priors from to-be-reasoned signals, enabling clean information ablations. Using DriveMind, we train representative VLM agents with Supervised Fine-Tuning and Group Relative Policy Optimization and evaluate them with nuPlan's metrics. Our results, unfortunately, indicate a consistent causal disconnect in reasoning-planning: removing ego/navigation priors causes large drops in planning scores, whereas removing CoT produces only minor changes. Attention analysis further shows that planning primarily focuses on priors rather than the CoT. Based on this evidence, we propose the Reasoning-Planning Decoupling Hypothesis, positing that the training-yielded reasoning is an ancillary byproduct rather than a causal mediator. To enable efficient diagnosis, we introduce a novel, training-free probe that measures an agent's reliance on priors by evaluating its planning robustness against minor input perturbations. In summary, we provide the community with a new dataset and a diagnostic tool to evaluate the causal fidelity of future models.
Figures & tables
| Agent | Base Model | Tuning | Vision | CoT | Priors |
| Supervised Fine-Tuning (SFT) Experiments | |||||
| Base | Qwen2.5-vl-7b | — | ✓ | — | — |
| CoT | Qwen2.5-vl-7b | SFT | ✓ | ✓ | All |
| Plan | Qwen2.5-vl-7b | SFT | ✓ | ✗ | All |
| Plan_NoV | Qwen2.5-vl-7b | SFT | ✗ | ✗ | All |
| CoT_NoHis | Qwen2.5-vl-7b | SFT | ✓ | ✓ | w/o History |
| Open-Loop Score | Closed-Loop Score | Avg. ADE / FDE (m) | Collision Ratio | |||||||||
| Driving Agent | 1s | 2s | 3s | 1s | 2s | 3s | 1s | 2s | 3s | 1s | 2s | 3s |
| Base | 36.52 | 33.62 | 31.50 | 59.80 | 42.48 | 36.99 | 2.67/4.45 | 4.80/8.81 | 7.02/13.75 | 7.84% | 9.16% | 9.87% |
| CoT | 98.91 | 96.67 | 92.56 | 97.46 | 96.42 | 92.38 | 0.03/0.08 | 0.14/0.39 | 0.33/1.02 | 0.00% | 0.29% | 2.15% |
| Omnidrive | 98.92 | 96.85 | 92.96 | 97.49 | 96.35 | 92.47 | 0.03/0.08 | 0.13/0.38 | 0.31/0.99 | 0.00% | 0.31% | 2.03% |
| Plan | 98.89 | 96.64 | 92.58 | 97.38 | 96.49 | 92.60 | 0.04/0.09 | 0.14/0.41 | 0.33/1.02 | 0.00% | 0.36% | 2.55% |
| Plan_NoV | 98.95 | 96.84 | 92.92 | 97.46 | 96.37 | 92.54 | 0.03/0.08 | 0.13/0.38 | 0.32/0.98 | 0.00% | 0.36% | 2.45% |
| Attention Target | Scenario Reasoning Task | Planning Task | ||||
| Shallow Layer | Middle Layer | Final Layer | Shallow Layer | Middle Layer | Final Layer | |
| Image Tokens (%) | 2.91 | 5.55 | 11.52 | 2.47 | 2.51 | 1.82 |
| Textual Priors (%) | 3.74 | 5.34 | 10.36 | 11.52 | 17.93 | 26.97 |
| Generated Reasoning (%) | — | — | — | 17.11 | 21.12 | 19.37 |
| All Textual Tokens (%) | 97.09 | 94.45 | 88.48 | 97.53 | 97.49 | 98.18 |
| 0.02 | 0.06 | 0.08 | 0.10 | 0.30 | |
| MFD | 0.55 | 2.17 | 3.75 | 5.89 | 15.46 |
| RID | 1.00 | 1.47 | 2.04 | 2.17 | 2.04 |
Appendix figures & tables12 assets
Supplementary material from the paper’s appendix.
Appendix
| Dataset | Real-World Logs | Semantic Richness | Expert CoT | Modular Structure | Scenario Coverage |
| DriveLM | Mixed | Moderate (nuScenes), High (CARLA) | Yes | No | — |
| DriveVLM | Yes | Moderate (SUP-AD) | Yes | No | 40 |
| DriveCoT | No | High (CARLA) | Yes | No | 5 |
| OmniDrive | Yes | Moderate (nuScenes) | No | No | — |
| DriveMind (Ours) | Yes | High (nuPlan) | Yes | Yes | 61 |
| Difficulty Category | Scenario Name | Sample Count | Total/Proportion |
| changing_lane_to_left | 56 | ||
| changing_lane_to_right | 51 | ||
| changing_lane_with_trail | 24 | ||
| crossed_by_vehicle | 41 | ||
| high_lateral_acceleration | 648 | ||
| starting_protected_cross_turn | 291 |
| Agent | Ablation Details with DriveMind |
| Base | Base model without any finetuning. Serves as a control group. |
| CoT | Finetuned on the full DriveMind dataset with 50K CoT VQA samples. |
| Plan | The reasoning part is removed from each sample in DriveMind, retaining only the final planning results. |
| Plan_NoV | Same as ‘Plan’, but all visual inputs are also removed. |
| CoT_NoHis | Ego’s history information is removed from the textual input of the samples in DriveMind. |
| CoT_NoHis_Ego | Both history and the ego state are removed from the textual input of the samples in DriveMind. |
| Attention Target | Detection Task | Planning Task | ||||
| Shallow Layer | Middle Layer | Final Layer | Shallow Layer | Middle Layer | Final Layer | |
| Image Tokens (%) | 3.46 | 10.02 | 15.92 | 1.96 | 3.22 | 4.05 |
| Textual Priors (%) | 6.76 | 5.93 | 6.89 | 10.52 | 9.93 | 23.97 |
| Textual Tokens (%) | 96.54 | 89.98 | 84.08 | 98.04 | 96.78 | 95.95 |
| Agent | Open-Loop Score | Close-Loop Score |
| CoT | 92.56 | 92.38 |
| CoT_V_Crop_Noise | 91.86 | 90.79 |
| CoT_V_Replace | 92.36 | 92.27 |
| Category | Scenario | Open-Loop Score | Closed-Loop Score | ||||
| CoT | CoT_NoPri | Plan_NoV | CoT | CoT_NoPri | Plan_NoV | ||
| Hard | high_lateral_acceleration | 89.69 | 40.74 | 90.04 | 90.31 | 66.18 | 90.51 |
| behind_long_vehicle | 98.31 | 67.25 | 98.57 | 99.28 | 77.74 | 99.45 | |
| changing_lane | 92.33 | 27.48 | 92.40 | 85.96 | 67.45 | 86.08 | |
| starting_left_turn | 86.98 | 36.94 | 87.26 | 82.93 | 69.91 | 83.41 | |
| starting_right_turn | 88.64 | 42.56 | 88.94 | 84.63 | 68.14 | 85.49 | |
| Agent | Open-Loop Score | Close-Loop Score |
| CoT_Seq | 91.64 | 91.45 |
| CoT_Seq_NoPri | 49.61 | 74.62 |
| Agent | MFD | RID |
| CoT_grpo_visual_match | 5.77 | 1.99 |
| Base_grpo_visual_match | 4.42 | 1.61 |