Vision-language-action (VLA) models struggle on history-dependent manipulation tasks, where the current observation alone does not determine the action, and the policy needs a memory of the history. Existing memory methods decide what to remember by design, for example, keeping frames with large pixel changes, and show inconsistent gains across tasks. We view what to remember as an optimisation problem. From the POMDP formulation of imitation learning, we show that the optimal memory maximises the conditional mutual information I(at;mt∣ot) between the action and the memory given the current observation. Intuitively, this means preserving the action-relevant information in the history that is not already contained in the current observation. Based on our analysis, we propose Divide-and-Remember (D&R), a recursive memory method that learns a memory function mt=M(ht) and scales to long contexts while staying compute-light. It involves two strategies: (1) the selection over the full history is divided recursively into subproblems of top-K selection over 2K tokens, so that fixed-size, lightweight selectors learned end-to-end support an unbounded history; (2) all recursion blocks share one selector, which captures the selection rule common to every block and keeps the method efficient. On RoboMME, a benchmark of 16 long-horizon manipulation tasks that require remembering when, where, what, and how to act, D&R achieves a state-of-the-art average success rate with consistent gains across all four suites under a budget of only 64 tokens; real-robot experiments show the same gain. Code, checkpoints and more results are at https://dnr-memory.github.io/
Figures & tables
Figure 1: Memory is a question of what to remember. (a) A history-dependent task: “watch the video carefully, then move the cube to the target in the same manner as before”. Two videos show different manners (pick-and-place versus peg push), yet at time step t the robot sees the same observation. A memoryless policy cannot recall how the cube was moved. (b) We view memory as an optimisation problem: find the memory mt=M(ht) , selected from the full history, that carries the key fact the current observation lacks (here, the grasp pose at t−3 ), so that the policy knows how to move the cube.
Figure 2: Graphical models of (a) imitation learning in a POMDP, where at depends on the history ht , and (b) a memory-augmented policy, where mt makes at independent of ht given (ot,mt) .
Figure 3: Overview of D&R. (a) At time step t , the selection over the full history Ht is divided recursively into subproblems of top- K selection over 2K tokens, each solved by the same lightweight selector fψ ; the kept tokens are merged pairwise and re-selected level by level, and at the root the selected tokens are passed to the policy through the attention mask. Tokens keep their temporal order from left to right at every level. (b) D&R and all baselines modulate the memory into the action expert through an adaptive LayerNorm, where action features cross-attend to the memory tokens.
Figure 4
Figure 5: Memory at step 422 of a MoveCube episode under a 64 -token budget, with the task goal “watch the video carefully, then move the cube to the target in the same manner as before”. FrameSamp keeps 4 evenly spaced frames. TokenDrop keeps the first frame and spends the rest of the budget on the patches whose pixels changed most. D&R selects 64 patches recursively, 35 of them from the video phase. Kept patches are shown, white ones are not in memory, and light blue borders mark the video phase. All tasks are in Appendix B.3.3 .
Motion-Centric
Time-Sensitive
Short-Horizon Video
Long-Horizon Video
Dynamic Scene-Change
Event-Salient
Type
Method
Pattern Lock
Route Stick
Swing Xtimes
Insert Peg
Avg
Stop Cube
Video Umsk
Video Repick
Video UmskS
Avg
Video PlcBtn
Video PlcOrd
Avg
Button Umsk
Button UmskS
Pick HighL
Avg
Pick Xtimes
Bin Fill
Move Cube
Avg
No memory
π0.5
2.89
4.67
35.56
1.56
11.17
6.67
20.44
0.44
18.67
13.18
31.11
25.78
28.45
22.22
6.67
11.33
13.41
42.89
30.00
26.00
32.96
Perceptual
FrameSamp
15.56
17.78
88.45
2.22
31.00
27.34
30.00
14.22
20.89
21.70
26.89
25.56
26.23
6.67
5.11
18.67
10.15
60.44
34.44
52.66
49.18
TokenDrop
9.56
8.89
36.67
2.44
14.39
5.78
28.67
1.56
16.67
15.63
26.67
19.56
23.12
16.67
6.44
14.00
12.37
39.78
27.78
24.44
30.67
Latent
HAMLET
7.33
9.56
81.11
1.11
24.78
39.33
27.78
20.67
25.11
24.52
27.78
21.78
24.78
26.67
21.78
26.00
24.82
84.67
34.89
60.00
59.85
RB-VLA
0.67
0.20
86.00
3.33
22.55
10.67
30.00
26.00
23.33
26.44
24.67
24.00
24.34
28.00
20.00
26.67
24.89
79.33
36.67
36.67
50.89
Table 1: Per-task success rate (%) on RoboMME ( Dai et al., 2026 ) , tasks grouped by functional characteristic (the same results by suite are in Table 7 ). All memories are built on π0.5 with a budget of 64 tokens and differ in what they store: raw patch tokens kept by frame sampling (FrameSamp) or token dropping (TokenDrop); learned moment tokens (HAMLET) or a recursive belief (RB-VLA); or fast weights (TTT). Each number is the mean over the last three checkpoints and three seeds; standard deviations are in Table 7 . Avg is the mean over the tasks of a characteristic; bold is the highest success rate in each column.
Figure 6: (a) Average success rate against the TFLOPs of one policy forward pass; each line sweeps the memory budget K from 64 to 512 (Table 8 ). (b) Average success rates against K . (c) D&R per task characteristic ( Dai et al., 2026 ) as the candidate pool ∣H∣ grows from 128 to 1024 tokens at K=64 ; mean ± std over runs (Table 5 ).
Method
Put Bottles
Track Cube
Repick Cube
Draw Pattern
Total
π0.5
0/10
2/10
0/10
1/10
3/40
FrameSamp
4/10
8/10
6/10
4/10
22/40
D&R
9/10
9/10
9/10
8/10
35/40
Table 2: Real-world results.
Appendix figures & tables27 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 7: Graphical models of imitation learning in a POMDP . (a) A general policy for the POMDP inputs the current observation and the full history, at∼π(⋅∣ot,ht) , but the history grows without bound. (b) A memoryless policy at∼π(⋅∣ot) inputs only the current observation, which causes the memory problem in Figure 1 (a). (c) A belief policy at∼π(⋅∣bt) inputs the belief bt(s)=P(st=s∣ht,ot) , the posterior over the current state and a sufficient statistic of the history. (d) A memory-augmented policy at∼π(⋅∣ot,mt) inputs a memory mt=M(ht) and the current observation. Here mt need only be sufficient to reproduce the expert’s action (Proposition 1 ), and the optimal memory carries the most action-relevant information I(at;mt∣ot) . Shaded circles are observed in the demonstrations, white circles are hidden.
Figure 8: Different methods assume different sufficiency metrics on the history. Bisimulation measures reward and transition; policy similarity measures policy and transition; the belief takes all three. D&R keeps policy sufficiency alone, under a limited budget.
Hyperparameter
Value
Base model
π0.5 base
Batch size
64 global
Action horizon
20
Action space
Joint-space
Proprioception
False
Total steps
80,000
Appendix
Table 3: Training hyperparameters.
Variant
Counting
Permanence
Reference
Imitation
AVG
∣H∣=128 (8 frames)
58.25 ± 2.04
20.83 ± 4.79
27.00 ± 1.38
34.67 ± 1.70
35.19 ± 1.98
∣H∣=256 (16 frames)
59.17 ± 2.84
25.17 ± 2.47
24.56 ± 2.13
38.67 ± 1.99
36.89 ± 0.81
∣H∣=512 (32 frames)
61.17 ± 0.62
26.83 ± 1.03
28.50 ± 1.41
37.83 ± 2.78
38.58 ± 0.50
∣H∣=1024 (64 frames)
61.89 ± 2.37
26.17 ± 2.32
20.94 ± 1.57
28.67 ± 2.58
34.42 ± 1.03
Appendix
Table 4: Candidate pool. Success rate (%) of D&R averaged over the four tasks of each suite, as mean ± standard deviation over evaluation runs (checkpoints × seeds 0/7/42 ), as the size ∣H∣ of the candidate pool the selector reduces grows from 8 to 64 frames. Every row uses a memory budget of K=64 tokens. bold is the highest success rate in each column.
Variant
Motion- Centric
Time- Sensitive
Short-Horizon Video
Long-Horizon Video
Dynamic Scene-Change
Event- Salient
AVG
∣H∣=128 (8 frames)
41.92 ± 1.62
20.67 ± 8.46
26.33 ± 1.58
33.17 ± 1.77
15.33 ± 5.84
61.11 ± 2.69
35.19 ± 1.98
∣H∣=256 (16 frames)
45.94 ± 2.05
28.00 ± 7.06
26.52 ± 1.80
30.67 ± 3.06
19.33 ± 1.99
59.85 ± 2.77
36.89 ± 0.81
∣H∣=512 (32 frames)
47.50 ± 1.08
31.33 ± 2.49
28.89 ± 1.37
31.67 ± 1.70
23.78 ± 1.26
58.22 ± 1.66
38.58 ± 0.50
∣H∣=1024 (64 frames)
36.17 ± 1.41
32.44 ± 5.56
23.19 ± 2.06
25.78 ± 2.15
22.44 ± 2.95
61.70 ± 3.02
34.42 ± 1.03
Appendix
Table 5: Candidate pool by task characteristic. Success rate (%) of D&R averaged over the tasks of each characteristic of Dai et al. (2026) (the grouping of Table 1 ), as mean ± standard deviation over evaluation runs (checkpoints × seeds 0/7/42 ), as the size ∣H∣ of the candidate pool changes. Every row uses a memory budget of K=64 tokens; the AVG column is that of Table 4 . bold is the highest success rate in each column.
Variant
Counting
Permanence
Reference
Imitation
AVG
D&R (reference)
59.17 ± 2.84
25.17 ± 2.47
24.56 ± 2.13
38.67 ± 1.99
36.89 ± 0.81
+ pred
58.62 ± 1.85
22.69 ± 1.52
22.12 ± 1.43
30.75 ± 2.96
33.55 ± 0.70
+ text
52.00 ± 2.31
21.94 ± 3.51
23.11 ± 2.02
24.22 ± 3.15
30.32 ± 1.54
Appendix
Table 6: Ablations. Success rate (%) of D&R and its variants at a candidate pool of ∣H∣=256 , as mean ± standard deviation over evaluation runs (checkpoints × seeds 0/7/42 ), averaged over the four tasks of each suite (top) and per task (bottom). + pred adds RB-VLA’s world-model loss to the action loss; + text adds the language instruction to the candidate pool the selector scores. Every row uses a memory budget of K=64 tokens. bold is the highest success rate in each column.
Type
Method
Counting
Permanence
Reference
Imitation
AVG
No memory
π0.5
28.78
17.00
17.17
8.78
17.93
Perceptual
FrameSamp
52.67
15.67
21.34
22.06
27.93
TokenDrop
27.50
17.11
15.45
11.33
17.85
Latent
HAMLET
60.00
25.34
24.06
19.50
32.22
RB-VLA
53.17
25.33
25.34
10.22
28.52
Parametric
TTT
28.84
16.72
13.61
9.34
17.13
Appendix
Table 7: Full results on RoboMME. Success rate (%) of every method under the 64 -token budget, grouped by what the memory stores. Top: mean over the four tasks of each suite and over all 16 tasks (AVG). Bottom: every task, as mean ± standard deviation over seeds. The π0.5 row is the number reported by Dai et al. (2026) . bold is the highest success rate in each column.
Overall
Budget
FrameSamp
TokenDrop
HAMLET
D&R
64
27.93
17.85
32.22
38.58
128
36.54
32.76
33.00
49.01
256
42.90
36.36
32.50
50.11
512
44.51
36.27
36.06
46.40
Appendix
Table 8: Effect of the memory budget. Average success rate (%) on RoboMME as the memory budget grows from 64 to 512 tokens, overall and per suite. The FrameSamp entries from 128 to 512 tokens and the TokenDrop entries at 128 and 256 tokens are the numbers reported by Dai et al. (2026) ; the remaining entries are our runs.
Counting
Permanence
Type
Method
BinFill
PickXtimes
SwingXtimes
StopCube
VideoUnmask
ButtonUnmask
VideoUnmaskSwap
ButtonUnmaskSwap
No memory
π0.5
30.00 ± 6.48
42.89 ± 6.41
35.56 ± 5.27
6.67 ± 5.20
20.44 ± 2.96
22.22 ± 10.22
18.67 ± 4.24
6.67 ± 4.12
Perceptual
FrameSamp
39.56 ± 5.27
87.33 ± 2.45
92.00 ± 2.24
42.00 ± 11.36
32.67 ± 3.16
25.11 ± 3.18
24.44 ± 6.06
18.22 ± 3.80
TokenDrop
31.33 ± 6.02
84.67 ± 4.84
86.00 ± 4.56
11.00 ± 4.69
30.67 ± 2.07
23.67 ± 4.46
22.67 ± 5.01
21.00 ± 6.03
Latent
HAMLET
31.78 ± 2.91
88.89 ± 3.48
89.33 ± 3.74
52.67 ± 7.42
29.78 ± 3.23
25.11 ± 4.26
23.33 ± 3.32
22.67 ± 3.61
RB-VLA
39.67 ± 3.13
78.31 ± 5.01
81.00 ± 8.50
12.02 ± 1.15
30.01 ± 3.21
24.20 ± 2.66
23.65 ± 0.13
19.10 ± 3.00
Appendix
Table 9: Results with a memory budget of 512 tokens. Per-task success rate (%) on RoboMME ( Dai et al., 2026 ) as mean ± standard deviation over seeds, with π0.5 as the backbone and methods grouped by what the memory stores. The π0.5 , FrameSamp, and TTT rows are the numbers reported by Dai et al. (2026) . bold is the highest success rate in each column.
Figure 9: BinFill (Counting): “put one red cube into the bin, then press the button to stop”.
Figure 10: PickXtimes (Counting): “pick up the green cube and place it on the target, then press the button to stop”.
Figure 11: SwingXtimes (Counting): “pick up the green cube, move it to the top of the right-side target, then move it to the top of the left-side target, repeating this back-and-forth motion two times, finally press the button to stop”.
Figure 12: StopCube (Counting): “press the button to stop the cube just as it reaches the target for the fourth time”.
Figure 13: VideoUnmask (Permanence): “watch the video carefully, then pick up the container hiding the green cube”.
Figure 14: ButtonUnmask (Permanence): “first press the button, then pick up the container hiding the green cube”.
Figure 15: VideoUnmaskSwap (Permanence): “watch the video carefully, then pick up the container hiding the green cube”.
Figure 16: ButtonUnmaskSwap (Permanence): “first press both buttons on the table, then pick up the container hiding the green cube”.
Figure 17: PickHighlight (Reference): “first press the button, then pick up all cubes that have been highlighted with white areas on the table”.
Figure 18: VideoRepick (Reference): “watch the video carefully, then pick up the same block that was previously picked up again, finally put it down and press the button to stop”.
Figure 19: VideoPlaceButton (Reference): “watch the video carefully, then place the green cube on the target right before the button was pressed”.
Figure 20: VideoPlaceOrder (Reference): “watch the video carefully, then place the green cube on the second target it was previously placed on”.
Figure 21: MoveCube (Imitation): “watch the video carefully, then move the cube to the target in the same manner as before”.
Figure 22: InsertPeg (Imitation): “watch the video carefully, then grasp the same end of the same peg you’ve picked before and insert it into the same side of the box”.
Figure 23: PatternLock (Imitation): “watch the video carefully, then use the stick attached to the robot to retrace the same pattern”.
Figure 24: RouteStick (Imitation): “watch the video carefully, then use the stick attached to the robot to navigate around the sticks on the table, following the same path”.
Figure 25: The four real-world tasks, recorded on our Franka Panda: PutBottles (Counting), TrackCube (Permanence), RepickCube (Reference) and DrawPattern (Imitation). For each task, the upper dashed box shows frames from the earlier history (the video the policy watched; for PutBottles, earlier frames of the episode, including a human intervention), and the lower box shows the front-camera observation during execution.
Figure 26: Real-robot platform. A 7-DoF Franka Emika Panda with a Robotiq 2F-140 gripper and two RealSense cameras: a wrist camera mounted on the gripper and a front camera on a tripod facing the workspace. The objects of the four tasks (baskets, cubes, cups, and bottles) are laid out on the table.
Vision-language-action (VLA) models have driven rapid progress in robotic manipulation, demonstrating strong fine-grained control and promising performance on long-horizon tasks. However, many existing VLAs lack explicit access to interaction history, making them vulnerable to perceptual aliasing: similar current observations and robot states at different task stages may induce action ambiguity and lower success rate. Existing methods incorporate temporal or progress cues through feature conditioning, action-prior modification, or sampling guidance. However, methods that jointly fine-tune memory modules and the base VLA incur additional policy-training costs, motivating the separation of trainable history-conditioned steering from frozen base-policy refinement. We propose ActMem-VLA, a dual-expert handover architecture that augments a frozen, fine-tuned VLA with a memory plugin comprising a Mamba-based memory module and a lightweight PreAction Expert (PAE). Specifically, Mamba encodes executed-action history into memory that conditions PAE alongside current context. With these inputs, PAE steers task progression during early, high-noise denoising, then passes the partially denoised action to the frozen Action Expert (AE) to refine action details during the remaining low-noise steps. The fine-tuned base VLA remains frozen throughout training, while only the Mamba module and PAE are jointly optimized. On LIBERO-Mem, ActMem-VLA achieves 80.8% average success across all ten tasks, compared with 65.2% for π0.5 and 49.5% for MemoryVLA, while introducing only 3.45% additional parameters. Across four real-world tasks, it improves the average success rate over π0.5 by 28.8%.
Yaxin Zhao, Dianye Huang, Chenwei Wang +2
Harbin Institute of Technology, Harbin, China. · Medical Intelligence and Robotic Cognition (MIRoC) Lab, Department of Mechanical Engineering, The University of Hong Kong (HKU), Hong Kong SAR, China. · Department of Computing, The Hong Kong Polytechnic University, Hong Kong SAR, China
Mainstream Vision-Language-Action (VLA) models predict actions primarily from the current observation under a Markovian assumption, thus struggling with long-horizon, temporally dependent tasks. Existing memory-augmented VLAs either expand the observation window or retrieve history from the memory bank as auxiliary policy-side context. However, they leave memory outside the native latent embedding space of VLA reasoning, preventing historical experience from being fluidly interleaved with multimodal reasoning and action formation. To this end, we introduce LaMem-VLA, a latent-memory-native framework that reconstructs historical experience into latent memory tokens and directly interweaves them with VLA reasoning. At its core, LaMem-VLA introduces four coordinated components: (i) a curator that organizes historical experience into two complementary short-term and long-term memory vaults; (ii) a seeker that queries both vaults using the multimodal cognition to retrieve context-relevant evidence; (iii) a condenser that reconstructs the retrieved evidence into compact short-term and long-term latent memory tokens; and (iv) a weaver that injects these memory tokens with the current observation and instruction into one continuous embedding sequence. By representing, retrieving, and consuming historical experience entirely in the same continuous latent space, LaMem-VLA enables memory to directly participate in VLA reasoning and guide action generation under a bounded context. Extensive experiments on SimplerEnv and LIBERO demonstrate the superiority of our LaMem-VLA.
Hongyu Qu, Jianzhe Gao, Xiaobin Hu +6
Nanjing University of Science and Technology · Zhejiang University · National University of Singapore
Vision-language-action policies often see only one or a few recent frames, which makes it difficult to evaluate how they use information that disappears during a task. We introduce MIKASA-Robo-VLA, a benchmark of 90 language-conditioned manipulation tasks. All but 10 hide the cue an action depends on. Those 10 are reactive controls. MIKASA-Robo, the suite it rebuilds, has 32 tasks and uses language only in a representative VLA subset. Here every task provides an instruction, while memory-dependent tasks hide a task-relevant cue and reactive controls keep it available. For 70 tasks, environment phase timings specify an information gap, and for 28 of them the gap exceeds the 16-frame window of the widest fixed-context VLA we survey. The gap counts only the interval the cue is provably absent, not the full duration a policy must retain it, so every memory-dependent task still requires memory by construction, including the ones whose measured gap is short. We release 22,500 oracle trajectories across 10 memory types in RLDS and LeRobotDataset v3. A reference π0.5 baseline with current images and proprioception, but no observation history or explicit memory module, is fine-tuned on 14 tasks and achieves 0.211 ± 0.044 mean task success. Its lower success on the evaluated Long-split tasks is confounded by open-loop chunking and the memory types represented in that subset. Project page: https://mikasarobo.github.io/
Egor Cherepanov, Nikita Kachaev, Aleksandr I. Panov +1
AXXX, Moscow, Russia · MIRIAI, Moscow, Russia · HSE University, Moscow, Russia