Sparse Feature Policy Unlearning Mitigates State Hallucination in Vision-Language-Action Models
Authors: Jiho Lee, Jeongeun Park, Heayoun Choi, Taekyung Kim, Eunwoo Kim
Organizations: School of Computer Science Engineering, Chung-Ang University, Seoul, 06974, Republic of Korea · School of Interactive Computing, Georgia Institute of Technology, Atlanta, GA, 30308, USA · NAVER AI Lab, Seongnam, 13561, Republic of Korea
Vision-Language-Action (VLA) models have shown strong generalization in robotic manipulation by leveraging rich representations from pretrained vision-language models. However, their deployment in real-world environments remains limited by recurring unreliable behaviors. In this work, we study state hallucination, a recurring failure pattern in which a VLA continues acting as if an unrealized robot-object state had been achieved. Our analyses find that state hallucination coincides with weakened attention to task-relevant visual regions, and a mechanistic interpretation via sparse autoencoders reveals that hallucination-associated sparse features are activated when these failures occur. Based on this analysis, we propose SOUL (Sparse feature pOlicy UnLearning), which selectively unlearns policy knowledge associated with state hallucination behaviors, where sparse features identified from hallucination failures and successful behaviors serve as explicit forgetting and retention targets, respectively. Experiments across VLA architectures in simulated and real-world environments show that our method substantially reduces hallucinated failures and improves task success without substantially compromising the existing manipulation capabilities. These results suggest that interpretable feature analysis provides a practical basis for selectively modifying undesirable knowledge in robot policies.
Figures & tables
Fig. 1 : Sparse feature policy unlearning mitigates hallucinated beliefs in VLAs. (a) Examples of grasp, transport, and place state hallucinations. (b) State hallucination occurs when a VLA mistakes its hallucinated belief for the actual robot-object state and generates actions based on this false belief. During this mismatch, hallucination features become strongly activated, and the proposed unlearning suppresses their activation while reinforcing success features. (c) Our sparse feature policy unlearning reduces hallucination while improving clean (hallucination-free) task success.
Fig. 2 : State hallucination is associated with distinct attention patterns. Representative clean success (CS, hallucination-free success) and hallucination failure (HF, failure with hallucination) rollouts under the same task show distinct attention patterns. HF exhibits weaker attention to task-relevant regions (e.g., gripper and target object) and increased attention to the background.
Fig. 3 : Sparse features associated with state hallucination can be identified. For each feature, we show clean success (CS) and hallucination failure (HF) rollouts with their temporal activations, ze,t,r,j , together with the distribution of rollout-level maximum activations, me,r,j . (a) A hallucination feature becomes strongly activated around hallucination events, while remaining weak in CS. (b) A success feature is strongly activated during successful subtask execution or task completion in CS, while remaining weak in HF. The boxplots show consistent activation differences between CS and HF across rollouts.
Fig. 4 : Overall framework. The proposed method provides a framework for improving robot policies by identifying and selectively modifying internal representations associated with undesirable behaviors. It first extracts visual-region-specific sparse features from policy activations and analyzes their activation patterns across execution outcomes, including clean success (CS), hallucination failure (HF), and non-hallucination failure (NF), characterized by task performance and hallucination occurrence. Based on these associations, region-feature pairs are selected as forget or retain targets. The policy is then updated to suppress hallucination-associated features while preserving reliable behavior. In this way, the policy is improved.
Method
Hallucination ↓
Hallucination failure ↓
Overall success ↑
Clean success ↑
Grasp Hall. ↓
Transport Hall. ↓
Place Hall. ↓
LIBERO-Plus with OpenVLA
Original Policy [ 2 ]
70.0
60.0
26.0
16.0
70.0
38.0
16.0
Gradient Ascent [ 31 ]
84.0
84.0
0.0
0.0
84.0
62.0
0.0
SOUL (Ours)
44.0
38.0
58.0
52.0
44.0
20.0
6.0
RoboCasa with π0.5
Original Policy [ 3 ]
40.0
37.1
41.9
39.0
39.0
7.6
1.9
TABLE I : Performance comparison of hallucination and task success on benchmarks. All values are reported as percentages (%). ‘Hall.’ denotes hallucination. ↓ and ↑ indicate that lower and higher values are better, respectively.
Fig. 5 : Visualizations of rollout behavior and feature activations on benchmarks. We compare the original policy and the proposed method, showing the temporal activations of selected forget and retain features during execution.
Fig. 6 : Real-world rollouts and selected sparse feature activations before and after applying the proposed unlearning, SOUL. We compare the original policy and SOUL across representative real-world rollouts spanning both models and tasks.
Fig. 7 : Behavior preservation and correction after applying SOUL. (a) Preservation measures the fraction of originally successful task executions and non-hallucinated events that remain successful and non-hallucinated, respectively, after SOUL. (b) Correction measures the fraction of originally failed task executions and hallucinated events that become successful and non-hallucinated, respectively, after SOUL. Higher values indicate better preservation or correction.
Method
Hall. ↓
Hall. failure ↓
Overall success ↑
Clean success ↑
Global
45.8
41.4
42.7
38.3
Tokenized
41.5
35.5
47.3
41.3
Region-guided
33.5
29.5
56.2
52.2
TABLE II : Comparison of feature selection at different spatial granularities. ‘Global’ selects features from activations pooled across all visual tokens, ‘Tokenized’ from individual visual-token activations, and ‘Region-guided’ from activations aggregated within each semantic region.
sr,jhall
sr,jagg
Lretain
Hall. ↓
Hall. failure ↓
Overall success ↑
Clean success ↑
✓
✓
✓
33.5
29.5
56.2
52.2
✓
✓
36.8
29.9
55.2
48.3
✓
✓
39.9
36.9
50.7
47.7
TABLE III : Ablation study of the forget score formulation and retain objective. The aggregated score sr,jagg=sr,jsep−sr,jsucc complements sr,jhall for forget feature selection.
Vision-Language-Action (VLA) policies translate language and visual inputs into robot actions, where their hidden representations directly shape closed-loop behavior. However, mechanistic interpretability tools from language and vision-language models do not transfer cleanly to VLAs: outputs are robot actions rather than human-readable tokens, and interventions can only be tested via expensive closed-loop rollouts. We propose an event-grounded interpretability pipeline that anchors SAE feature analysis to behavioral events rather than text contexts. End-effector keyframes are clustered within each task using visual, state, and temporal cues, linking SAE features to behaviorally salient events and, via optional VLM annotations, to semantic context. To our knowledge, our pipeline is among the first to ground SAE-based VLA analysis in closed-loop behavioral events. Across two simulation architectures and a real-robot study, event-grounded ranking yields the strongest causal effects on OpenVLA and transfers to the continuous action chunks of π0.5. SAE is a sparse but imperfect intervention basis: usability varies with architecture and intervention site, and aggressive intervention reveals safety and interpretability limits. Overall, event-grounded SAE analysis emerges as a practical starting point for behavior-anchored VLA interpretability, motivating future work on SAE features beyond action-aligned coordinates, finer-grained closed-loop evaluation, and safe interventions for high-stakes VLA deployments. Code is available at https://github.com/xc-j/Event-SAE.
Xinchen Jin, Aditya Chatterjee, Pranav Kumar +1
Department of Computer Science, Purdue University West Lafayette, IN 47907
Recent Vision-Language-Action (VLA) models achieve promising performance in robotic manipulation, typically measured by success rates aggregated over predefined object configurations, an evaluation that implicitly assumes spatially uniform competence across the workspace. However, this assumption does not hold: even with the instruction and every other scene factor held fixed, merely relocating a task-irrelevant distractor can sharply raise the failure probability within localized, spatially coherent regions, which we term Positional Blind Spots (PBS). In this paper, we propose a two-stage black-box framework to uncover and mitigate PBS. During the uncovering stage, we grid the workspace and apply a one-sided log-likelihood-ratio test to localize PBS cells with significantly elevated risk. During the mitigation stage, we fine-tune the policy via LoRA on demonstrations collected from these PBS regions, improving competence there while largely preserving performance across the rest of the workspace. We evaluate our framework on five state-of-the-art VLA policies across two benchmarks, and find that PBS are pervasive and spatially concentrated in all of them, with failure rates up to 0.58. Our search strategy achieves an average F1-score of 0.678, outperforming random search and adaptive sampling baselines by 0.268 and 0.178, respectively. Guided by the discovered regions, targeted fine-tuning reduces the overall failure rate by 40.00%--85.19%.
Dongdong An, Pengjie Zhao, Yihao Huang +5
Shanghai Normal University Shanghai, China · East China Normal University Shanghai, China · Northwest A&F University Yangling, Shaanxi, China +1
Vision-Language-Action models face significant challenges in real-world deployment due to the entanglement of high-level reasoning with low-level control, and the instability of policy optimization. In this paper, we introduce SyVLA, a robust VLA model trained with diversified experiences. We propose an Intention Decoupling algorithm to isolate control-relevant features from reasoning contexts and a similar-sample guided RL pipeline to stabilize policy updates and mitigate distribution shift. Extensive experiments on real-world robotic tasks and multi-modal benchmarks demonstrate that SyVLA achieves superior task success rates and stronger out-of-distribution generalization compared to existing methods, while effectively preserving core vision-language capabilities. Codes and Datasets is released on project page.
Leiyu Wang, Zhaofengnian Wang, Xueqi Li +3
Shanghai Jiao Tong University · Shanghai Innovation Institute · Tongji University +3