Sparse Feature Policy Unlearning Mitigates State Hallucination in Vision-Language-Action Models
Authors: Jiho Lee, Jeongeun Park, Heayoun Choi, Taekyung Kim, Eunwoo Kim
Organizations: School of Computer Science Engineering, Chung-Ang University, Seoul, 06974, Republic of Korea · School of Interactive Computing, Georgia Institute of Technology, Atlanta, GA, 30308, USA · NAVER AI Lab, Seongnam, 13561, Republic of Korea
Vision-Language-Action (VLA) models have shown strong generalization in robotic manipulation by leveraging rich representations from pretrained vision-language models. However, their deployment in real-world environments remains limited by recurring unreliable behaviors. In this work, we study state hallucination, a recurring failure pattern in which a VLA continues acting as if an unrealized robot-object state had been achieved. Our analyses find that state hallucination coincides with weakened attention to task-relevant visual regions, and a mechanistic interpretation via sparse autoencoders reveals that hallucination-associated sparse features are activated when these failures occur. Based on this analysis, we propose SOUL (Sparse feature pOlicy UnLearning), which selectively unlearns policy knowledge associated with state hallucination behaviors, where sparse features identified from hallucination failures and successful behaviors serve as explicit forgetting and retention targets, respectively. Experiments across VLA architectures in simulated and real-world environments show that our method substantially reduces hallucinated failures and improves task success without substantially compromising the existing manipulation capabilities. These results suggest that interpretable feature analysis provides a practical basis for selectively modifying undesirable knowledge in robot policies.
Figures & tables
Fig. 1 : Sparse feature policy unlearning mitigates hallucinated beliefs in VLAs. (a) Examples of grasp, transport, and place state hallucinations. (b) State hallucination occurs when a VLA mistakes its hallucinated belief for the actual robot-object state and generates actions based on this false belief. During this mismatch, hallucination features become strongly activated, and the proposed unlearning suppresses their activation while reinforcing success features. (c) Our sparse feature policy unlearning reduces hallucination while improving clean (hallucination-free) task success.
Fig. 2 : State hallucination is associated with distinct attention patterns. Representative clean success (CS, hallucination-free success) and hallucination failure (HF, failure with hallucination) rollouts under the same task show distinct attention patterns. HF exhibits weaker attention to task-relevant regions (e.g., gripper and target object) and increased attention to the background.
Fig. 3 : Sparse features associated with state hallucination can be identified. For each feature, we show clean success (CS) and hallucination failure (HF) rollouts with their temporal activations, ze,t,r,j , together with the distribution of rollout-level maximum activations, me,r,j . (a) A hallucination feature becomes strongly activated around hallucination events, while remaining weak in CS. (b) A success feature is strongly activated during successful subtask execution or task completion in CS, while remaining weak in HF. The boxplots show consistent activation differences between CS and HF across rollouts.
Fig. 4 : Overall framework. The proposed method provides a framework for improving robot policies by identifying and selectively modifying internal representations associated with undesirable behaviors. It first extracts visual-region-specific sparse features from policy activations and analyzes their activation patterns across execution outcomes, including clean success (CS), hallucination failure (HF), and non-hallucination failure (NF), characterized by task performance and hallucination occurrence. Based on these associations, region-feature pairs are selected as forget or retain targets. The policy is then updated to suppress hallucination-associated features while preserving reliable behavior. In this way, the policy is improved.
Method
Hallucination ↓
Hallucination failure ↓
Overall success ↑
Clean success ↑
Grasp Hall. ↓
Transport Hall. ↓
Place Hall. ↓
LIBERO-Plus with OpenVLA
Original Policy [ 2 ]
70.0
60.0
26.0
16.0
70.0
38.0
16.0
Gradient Ascent [ 31 ]
84.0
84.0
0.0
0.0
84.0
62.0
0.0
SOUL (Ours)
44.0
38.0
58.0
52.0
44.0
20.0
6.0
RoboCasa with π0.5
Original Policy [ 3 ]
40.0
37.1
41.9
39.0
39.0
7.6
1.9
TABLE I : Performance comparison of hallucination and task success on benchmarks. All values are reported as percentages (%). ‘Hall.’ denotes hallucination. ↓ and ↑ indicate that lower and higher values are better, respectively.
Fig. 5 : Visualizations of rollout behavior and feature activations on benchmarks. We compare the original policy and the proposed method, showing the temporal activations of selected forget and retain features during execution.
Fig. 6 : Real-world rollouts and selected sparse feature activations before and after applying the proposed unlearning, SOUL. We compare the original policy and SOUL across representative real-world rollouts spanning both models and tasks.
Fig. 7 : Behavior preservation and correction after applying SOUL. (a) Preservation measures the fraction of originally successful task executions and non-hallucinated events that remain successful and non-hallucinated, respectively, after SOUL. (b) Correction measures the fraction of originally failed task executions and hallucinated events that become successful and non-hallucinated, respectively, after SOUL. Higher values indicate better preservation or correction.
Method
Hall. ↓
Hall. failure ↓
Overall success ↑
Clean success ↑
Global
45.8
41.4
42.7
38.3
Tokenized
41.5
35.5
47.3
41.3
Region-guided
33.5
29.5
56.2
52.2
TABLE II : Comparison of feature selection at different spatial granularities. ‘Global’ selects features from activations pooled across all visual tokens, ‘Tokenized’ from individual visual-token activations, and ‘Region-guided’ from activations aggregated within each semantic region.
sr,jhall
sr,jagg
Lretain
Hall. ↓
Hall. failure ↓
Overall success ↑
Clean success ↑
✓
✓
✓
33.5
29.5
56.2
52.2
✓
✓
36.8
29.9
55.2
48.3
✓
✓
39.9
36.9
50.7
47.7
TABLE III : Ablation study of the forget score formulation and retain objective. The aggregated score sr,jagg=sr,jsep−sr,jsucc complements sr,jhall for forget feature selection.