Remember What You Did: Action-History Memory with Dual-Expert Denoising for Long-Horizon Vision-Language-Action Policies
Organizations: Harbin Institute of Technology, Harbin, China. · Medical Intelligence and Robotic Cognition (MIRoC) Lab, Department of Mechanical Engineering, The University of Hong Kong (HKU), Hong Kong SAR, China. · Department of Computing, The Hong Kong Polytechnic University, Hong Kong SAR, China
Abstract
Vision-language-action (VLA) models have driven rapid progress in robotic manipulation, demonstrating strong fine-grained control and promising performance on long-horizon tasks. However, many existing VLAs lack explicit access to interaction history, making them vulnerable to perceptual aliasing: similar current observations and robot states at different task stages may induce action ambiguity and lower success rate. Existing methods incorporate temporal or progress cues through feature conditioning, action-prior modification, or sampling guidance. However, methods that jointly fine-tune memory modules and the base VLA incur additional policy-training costs, motivating the separation of trainable history-conditioned steering from frozen base-policy refinement. We propose ActMem-VLA, a dual-expert handover architecture that augments a frozen, fine-tuned VLA with a memory plugin comprising a Mamba-based memory module and a lightweight PreAction Expert (PAE). Specifically, Mamba encodes executed-action history into memory that conditions PAE alongside current context. With these inputs, PAE steers task progression during early, high-noise denoising, then passes the partially denoised action to the frozen Action Expert (AE) to refine action details during the remaining low-noise steps. The fine-tuned base VLA remains frozen throughout training, while only the Mamba module and PAE are jointly optimized. On LIBERO-Mem, ActMem-VLA achieves 80.8% average success across all ten tasks, compared with 65.2% for and 49.5% for MemoryVLA, while introducing only 3.45% additional parameters. Across four real-world tasks, it improves the average success rate over by 28.8%.
Figures & tables
| Variant | T6 | T7 | T8 | Avg. |
|---|---|---|---|---|
| w/o Execution Memory | 38.3 | 33.3 | 50.0 | 40.6 |
| w/o PreAction Expert | 58.3 | 38.3 | 63.3 | 53.3 |
| ActMem-VLA | 80.0 | 40.0 | 45.0 | 55.0 |
| Parameter | Default Tested | T6 | T7 | T8 | Avg. |
|---|---|---|---|---|---|
| Handover point | 65.0 | 36.7 | 43.3 | 48.3 | |
| 66.7 | 25.0 | 43.3 | 45.0 | ||
| 71.7 | 28.3 | 41.7 | 47.2 | ||
| Mamba depth | 58.3 | 25.0 | 46.7 | 43.3 | |
| Mamba width | 76.7 | 40.0 | 50.0 | 55.6 | |
| PAE depth | 70.0 | 21.7 | 40.0 | 43.9 |