cs.ROOct 7, 2026

Sparse Feature Policy Unlearning Mitigates State Hallucination in Vision-Language-Action Models

Authors: Jiho Lee, Jeongeun Park, Heayoun Choi, Taekyung Kim, Eunwoo Kim

Organizations: School of Computer Science Engineering, Chung-Ang University, Seoul, 06974, Republic of Korea · School of Interactive Computing, Georgia Institute of Technology, Atlanta, GA, 30308, USA · NAVER AI Lab, Seongnam, 13561, Republic of Korea

Abstract

Vision-Language-Action (VLA) models have shown strong generalization in robotic manipulation by leveraging rich representations from pretrained vision-language models. However, their deployment in real-world environments remains limited by recurring unreliable behaviors. In this work, we study state hallucination, a recurring failure pattern in which a VLA continues acting as if an unrealized robot-object state had been achieved. Our analyses find that state hallucination coincides with weakened attention to task-relevant visual regions, and a mechanistic interpretation via sparse autoencoders reveals that hallucination-associated sparse features are activated when these failures occur. Based on this analysis, we propose SOUL (Sparse feature pOlicy UnLearning), which selectively unlearns policy knowledge associated with state hallucination behaviors, where sparse features identified from hallucination failures and successful behaviors serve as explicit forgetting and retention targets, respectively. Experiments across VLA architectures in simulated and real-world environments show that our method substantially reduces hallucinated failures and improves task success without substantially compromising the existing manipulation capabilities. These results suggest that interpretable feature analysis provides a practical basis for selectively modifying undesirable knowledge in robot policies.

Figures & tables

Explore similar work

CardsList
  1. Event-Grounded Sparse Autoencoders for Vision-Language-Action Policies

    May 17, 2026Xinchen Jin, Aditya Chatterjee, Pranav Kumar +1VLM InterpretabilitySparse Autoencoders

  2. Uncovering and Mitigating Positional Blind Spots in Vision-Language-Action Models

    Aug 3, 2026Dongdong An, Pengjie Zhao, Yihao Huang +5Vision-Language-Action ModelsVLMs for Robotics

  3. Scaling by Diversified Experience for Vision-Language-Action Models

    Jun 8, 2026Leiyu Wang, Zhaofengnian Wang, Xueqi Li +3VLM RobustnessVision-Language-Action Models