Long-horizon manipulation requires robots to remember cues that are no longer in view while responding to moving objects. Yet vision-language-action (VLA) policies often rely on the latest observation, and refreshing their visual context typically requires another costly vision-language model (VLM) pass. We present D2-VLA, which combines dual memory and dual-frequency control at the KV-cache interface of a pretrained VLA. D2-VLA uses block-wise causal KV caching to encode observations incrementally and, guided by distinct temporal attention patterns, constructs separate historical KV read views for the VLM and action expert. Between periodic VLM updates, a gated adapter incorporates fresh visual features into the latest history-conditioned KV block, while a short fast-memory queue supports action replanning. We introduce DOMINO-Long, a ten-task benchmark requiring robots to use earlier visual cues when manipulating moving objects. D2-VLA achieves complete-task success rates of 29.3% on DOMINO, compared with 9.6% for π0.5 and 17.2% for PUMA, and 60.0% on DOMINO-Long, compared with 35.4% and 20.6%, respectively. It improves success rates on eight real-robot tasks and reaches 97.5% on LIBERO-Long and 74.3% on RoboTwin 2.0.
Figures & tables
Figure 1: Persistent memory and responsive control with D 2 -VLA. Unlike single-frame π0.5 (top), D 2 -VLA (bottom) retains temporal context through causal KV caching and incorporates fresh observations through a KV adapter for fast replanning between slow VLM updates. Real-world examples and benchmark success rates illustrate improvements in long-horizon dynamic manipulation.
Figure 2: Overview of D 2 -VLA. (a) Dual memory provides component-specific KV reads; dual-frequency control combines slow VLM refreshes with fast action replanning. (b) Block-causal attention allows bidirectional reads within each observation block. (c) The adapter refines the latest slow KV block with current visual features.
Figure 3: Attention asymmetry and token concentration. Visual attention maps (left) and token-wise attention scores (right) for the VLM (top) and action expert (bottom).
Method
SR (%) ↑
MS ↑
OpenVLA ( Kim et al., 2024 )
1.5
6.1
RDT-1B ( Liu et al., 2025 )
5.3
17.7
π0 ( Black et al., 2026 )
8.2
24.0
π0.5 ( Intelligence et al., 2025 )
9.6
26.2
InternVLA-M1 ( Chen et al., 2025b )
5.4
27.6
VLA-Adapter ( Wang et al., 2025 )
4.4
24.3
Table 1: Evaluation on dynamic and long-horizon manipulation. (a) Comparison on DOMINO, measured by SR and MS. (b) Per-subtask results on the 10 subtasks of DOMINO-Long.
Figure 4: Long-horizon dynamic real-world evaluation on the Unitree G1D robot. Representative task executions and success-rate comparisons with π0.5 and MemoryVLA.
Method
Params (B) ↓
SR (%) ↑
Diffusion Policy ( Chi et al., 2024 )
0.1
28.0
ACT ( Zhao et al., 2023 )
0.1
29.7
DP3 ( Ze et al., 2024 )
0.3
55.2
π0 ( Black et al., 2026 )
3.2
46.4
FlowPolicy ( Zhang et al., 2024 )
0.3
41.0
RDT-1B ( Liu et al., 2025 )
1.7
34.5
Table 2: Performance comparison on static manipulation benchmarks. (a) Model size and average success rate on RoboTwin 2.0, with all methods trained and evaluated exclusively on the clean setting. (b) Average success rate on LIBERO-Long.
Model
DM
DF
SR (%)
π0.5
–
–
9.6
+ Causal Training
✓
–
16.1
D 2 -VLA
✓
✓
29.3
Table 3: Component ablation on DOMINO. DM: dual-memory; DF: dual-frequency.
Appendix figures & tables12 assets
Supplementary material from the paper’s appendix.
Appendix
Consumer
Source
Current
Ret.
VLM prefill
Slow
–
0.7
Slow denoising
Slow
Yes
0.7
Fast denoising
Fast FIFO
Yes
0.7
ρ
0.2 (fixed)
Appendix
Table 4: Online KV selection settings for VLM prefill, slow-branch denoising, and fast-branch denoising. Source, Current, and Ret. denote historical source, current-block protection, and visual-token retention ratio, respectively.
Setting
Value
Initialization / training
π0.5 Base / full fine-tuning
Slow / fast cache capacity
4/2 complete blocks
Action horizon H / executed chunk size E
25/25
Slow-update period
V=3E
Main-result historical visual retention
0.7 for all active consumers
Supervised branches
3
Appendix
Table 5: Training and inference settings for DOMINO, DOMINO-Long, and the real-world tasks. Training steps count optimizer updates. Capacities include the current block.
Task
π0.5
PUMA
D 2 -VLA
SR (%)↑
MS (%)↑
SR (%)↑
MS (%)↑
SR (%)↑
MS (%)↑
Adjust Bottle
52.00
71.45
65.00
74.04
75.00
78.18
Beat Block Hammer
10.00
23.22
15.00
29.09
23.00
32.29
Click Alarmclock
14.00
18.47
4.00
8.38
5.00
8.74
Click Bell
0.00
6.35
3.00
13.64
11.00
19.00
Dump Bin Bigbin
0.00
16.64
0.00
39.05
38.00
44.32
Appendix
Table 6: Per-task DOMINO results under L1 + clean evaluation. Baseline SR (%) and MS for π0.5 and PUMA are transcribed from Table 16 of the DOMINO paper; D 2 -VLA results use a historical visual KV retention ratio of 0.7. All scores are displayed to two decimal places.
Figure 5: Overview of DOMINO-Long tasks. Visualization of the ten long-horizon manipulation tasks introduced in DOMINO-Long.
ID
Task
Information retained for later action
1
Bottle Return
Original region of the moving bottle.
2
Can Grasp
Identity of the previously observed can.
3
Stack Two
Presentation order of two blocks.
4
Two Color Cue
Colors specifying the object and destination.
5
Can Placement
Basket indicated by a transient visual cue.
6
Tray Swap
Original source region of the transferred block.
Appendix
Table 7: DOMINO-Long tasks and their memory requirements.
Task
SR
Task
SR
Task
SR
Adjust Bottle
100%
Open Microwave
86%
Place Object Stand
90%
Beat Block Hammer
82%
Pick Diverse Bottles
48%
Place Phone Stand
72%
Blocks Ranking RGB
72%
Pick Dual Bottles
62%
Place Shoe
82%
Blocks Ranking Size
64%
Place A2B Left
70%
Press Stapler
86%
Click Alarmclock
100%
Place A2B Right
72%
Put Bottles Dustbin
68%
Click Bell
100%
Place Bread Basket
58%
Put Object Cabinet
38%
Appendix
Table 8: Complete-task success rate (SR, %) on all 50 RoboTwin 2.0 Chen et al. (2025a) tasks. Task names are formatted from the corresponding environment identifiers.
Figure 6: Long-horizon tasks.
Figure 7: Long-horizon dynamic tasks.
Task
π0.5
M-VLA
Ours
L1: Ordered Cup Stacking
70
60
80
L2: Cross-Plate Object Swap
50
40
75
L3: Ordered Object Retrieval
45
35
65
L4: Original-Layout Restoration
40
25
65
Average
51.25
40
71.25
Appendix
Table 9: Complete-task SR (%) on the four long-horizon static real-world tasks. M-VLA and Ours denotes Memory-VLA and D2−VLA , respectively.
Figure 8: Historical KV-cache activation distributions on LIBERO and RoboTwin. In each subfigure, the left and right heatmaps show the magnitudes of cached keys (K) and values (V), respectively. The distributions are non-uniform, with stronger activations concentrated in a subset of tokens and feature dimensions. This concentration motivates selective KV-cache compression that prioritizes salient historical entries, suggesting the potential to reduce memory and computation with limited performance loss.
Figure 9: Accuracy–efficiency and historical KV-retention analysis. (a) Comparison with the single-observation baseline and Full KV Cache. LBO-L, RBT2.0, and DMO denote the LIBERO-Long Dai et al. (2026) , RoboTwin 2.0 Chen et al. (2025a) , and DOMINO Dai et al. (2026) benchmarks, respectively. SR, Lat., and Mem. refer to average success rate, inference latency, and peak GPU memory usage; (b) DOMINO success rate under different historical visual KV-retention ratios.
Setting
Period ratio V/E
SR (%)
Without dual-frequency mechanism
–
16.1
V:E=50:25
2
27.6
V:E=75:25
3
29.3
Appendix
Table 10: Dual-frequency ablation on DOMINO. The enabled settings use E=25 and vary the slow-update period V .
Vision-language-action (VLA) models have driven rapid progress in robotic manipulation, demonstrating strong fine-grained control and promising performance on long-horizon tasks. However, many existing VLAs lack explicit access to interaction history, making them vulnerable to perceptual aliasing: similar current observations and robot states at different task stages may induce action ambiguity and lower success rate. Existing methods incorporate temporal or progress cues through feature conditioning, action-prior modification, or sampling guidance. However, methods that jointly fine-tune memory modules and the base VLA incur additional policy-training costs, motivating the separation of trainable history-conditioned steering from frozen base-policy refinement. We propose ActMem-VLA, a dual-expert handover architecture that augments a frozen, fine-tuned VLA with a memory plugin comprising a Mamba-based memory module and a lightweight PreAction Expert (PAE). Specifically, Mamba encodes executed-action history into memory that conditions PAE alongside current context. With these inputs, PAE steers task progression during early, high-noise denoising, then passes the partially denoised action to the frozen Action Expert (AE) to refine action details during the remaining low-noise steps. The fine-tuned base VLA remains frozen throughout training, while only the Mamba module and PAE are jointly optimized. On LIBERO-Mem, ActMem-VLA achieves 80.8% average success across all ten tasks, compared with 65.2% for π0.5 and 49.5% for MemoryVLA, while introducing only 3.45% additional parameters. Across four real-world tasks, it improves the average success rate over π0.5 by 28.8%.
Yaxin Zhao, Dianye Huang, Chenwei Wang +2
Harbin Institute of Technology, Harbin, China. · Medical Intelligence and Robotic Cognition (MIRoC) Lab, Department of Mechanical Engineering, The University of Hong Kong (HKU), Hong Kong SAR, China. · Department of Computing, The Hong Kong Polytechnic University, Hong Kong SAR, China
Vision-language-action (VLA) models increasingly condition robot policies on history, depth, or 4D features to resolve ambiguity in long-horizon manipulation. However, more spatiotemporal evidence is not necessarily better: when the injected evidence is not motion-consistent, it can introduce geometric drift, fragmented temporal cues, and unstable action generation. This raises a simple question: should a VLA remember past frames, or remember the motion that connects them? We introduce MotionVLA, a motion-history interface that converts a short past-only video window into compact, time-continuous trajectory-field tokens. Instead of treating history as a sparse set of ndependently lifted frames, MotionVLA represents recent observations as physically coherent motion evidence. Current visual tokens query this history to retrieve task-relevant motion information, which is then recoupled into the VLA stream under trajectory-grounded supervision. Experiments across simulation benchmarks and preliminary real-robot rollouts show that MotionVLA improves long-horizon manipulation while producing smoother and more direct executions. These results suggest that effective VLA memory is not just about providing more 4D context, but about exposing motion-consistent evidence that is usable for control.
Shanglin Yuan, Weiheng Zhao, Xianda Guo +4
Huazhong University of Science and Technology · D-Robotics · Wuhan University
Mainstream Vision-Language-Action (VLA) models predict actions primarily from the current observation under a Markovian assumption, thus struggling with long-horizon, temporally dependent tasks. Existing memory-augmented VLAs either expand the observation window or retrieve history from the memory bank as auxiliary policy-side context. However, they leave memory outside the native latent embedding space of VLA reasoning, preventing historical experience from being fluidly interleaved with multimodal reasoning and action formation. To this end, we introduce LaMem-VLA, a latent-memory-native framework that reconstructs historical experience into latent memory tokens and directly interweaves them with VLA reasoning. At its core, LaMem-VLA introduces four coordinated components: (i) a curator that organizes historical experience into two complementary short-term and long-term memory vaults; (ii) a seeker that queries both vaults using the multimodal cognition to retrieve context-relevant evidence; (iii) a condenser that reconstructs the retrieved evidence into compact short-term and long-term latent memory tokens; and (iv) a weaver that injects these memory tokens with the current observation and instruction into one continuous embedding sequence. By representing, retrieving, and consuming historical experience entirely in the same continuous latent space, LaMem-VLA enables memory to directly participate in VLA reasoning and guide action generation under a bounded context. Extensive experiments on SimplerEnv and LIBERO demonstrate the superiority of our LaMem-VLA.
Hongyu Qu, Jianzhe Gao, Xiaobin Hu +6
Nanjing University of Science and Technology · Zhejiang University · National University of Singapore