ReCAT: Remember, Count, and Time: Structured Recurrent Memory for Robot Manipulation
Authors: Pankhuri Vanjani, Mostafa Hatab, Can Mizrakli, Vaisakh Shaj, Zhuoyue Li, Moritz Reuss, Rudolf Lioutikov
Organizations: Intuitive Robots Lab, Karlsruhe Institute of Technology (KIT), Germany · University of Edinburgh, UK · NVIDIA · Robotics Institute Germany (RIG)
Memory-dependent manipulation requires robots to make decisions using information that is no longer available to their current sensors, such as recalling an earlier visual cue, tracking task progress, counting repeated events, or estimating elapsed time. We present ReCAT, a language-conditioned policy with structured recurrent memory. An instruction-conditioned encoder forms features from the current observation. A recurrent memory integrates the observation stream through Mamba-2 layers and one causal attention layer. A flow-matching Transformer decoder reads the current and the historical representation through separate cross-attention in every block. ReCAT reaches 95.3% average success on LIBERO and 62.4% on RMBench, with the best or tied-best result on six of nine tasks. On three real-robot tasks probing spatial recall, event counting, and interval timing, the best ReCAT variant reaches 66.7% average success, against 8.3% for the strongest short-history baseline. Controlled comparisons within ReCAT show that the observation encoder and every-block memory conditioning are needed for this performance. They also show that update rules developed for efficient sequence modeling behave differently as robot memory: additive updates have the highest observed success on counting and timing, and delta-rule updates on spatial recall. Project website is at https://intuitive-robots.github.io/ReCAT
Figures & tables
Fig. 1: Visually similar observations can require different actions depending on task history. ReCAT retains the completed scoop count in a compact recurrent memory, allowing the robot to scoop again after one scoop and put down the shovel after two.
Fig. 2: Overview of ReCAT. Frozen DynaFLIP and T5 encoders (LoRA-adapted) feed a language-conditioned visual compressor and a six-layer Mamba–attention frame encoder, which produces one feature vector zt per step. The same zt feeds the action decoder directly and is the input to the recurrent memory, which returns the readout mt . The flow-matching decoder attends to zt and mt through separate cross-attention in every block. The decoder backbone is shared across all experiments. The ablations vary the frame-encoder mixer, the recurrent update rule, memory depth and width, and how mt enters the decoder.
Model
State update
Update semantics
Mamba-2
diag(at)ht−1+vtkt⊤
Accumulate: storing a similar cue again adds another write to its trace, so repeated occurrences of an event build up in the state, the property counting requires.
Mamba-3
diag(ateiθt)ht−1+vtkt⊤
Accumulate + rotate: as above, with phase evolution that provides an additional signal for temporal progression or motion.
GDN-2
αt(I−βtktkt⊤)ht−1+βtvtkt⊤
Replace: storing a similar cue overwrites its previous content with the latest value, while other stored cues remain undisturbed. Repeated identical events therefore leave a single trace.
TABLE I: Update rules of the recurrent layers compared in Sec. IV . Each step writes a cue-content pair (kt,vt) derived from zt into the state ht . The rules differ in what the write does to information already stored. The descriptions give the intended semantics of each rule. Whether a trained policy uses them this way is an empirical question.
Fig. 3: Real-robot tasks, one per row, each isolating one memory demand: plant counts scoops across five stages, pot timer times an interval while the scene is static, and sponge recalls an unmarked start position that has left the camera view. Outlined frames mark the decision point where the correct action depends on an earlier event, not on the current observation.
Method
Spatial
Object
Goal
LIBERO-10
Avg.
LIBERO-Plus
Diffusion policy
78.3
92.5
68.3
50.5
72.4
–
X-VLA [ 42 ]
98.2
98.6
97.8
97.6
98.1
71.4
Ours
96.4
97.2
97.2
90.4
95.3
64.2
TABLE II: Success rates (%) on LIBERO.
Task
TMC
DP (90M)
ACT (80M)
\/∗pi0.5 (3B)
*X-VLA (0.9B)
MEM-0 (10B)
EventVLA (4B)
*MemoryVLA (7.3B)
*MemER (10B)
Ours (374M)
Observe and Pick Up
M(1)
1
1
9
9
4
21
2
7
4
Rearrange Blocks
M(1)
0
29
13
13
89
96
53
17
96
Put Back Block
M(1)
0
0
11
18
90
95
81
0
100
Swap Blocks
M(1)
11
2
24
16
67
96
76
14
100
Swap T
M(1)
20
2
15
3
14
87
9
7
64
Average
M(1)
6.4
6.8
14.4
11.8
52.8
79.0
44.2
9.0
72.8
TABLE III: Success rates (%) on RMBench. * Indicates models utilizing large scale pretraining.
Plant
Pot timer
Sponge
Method
1 scoop
2 scoops
3 or more
Flower planted Success
Pot on stove
No wait
Wrong wait
≈ 30 s wait
Pepper in the pot
Success
Pick sponge, put on plate
Returned to wrong position
Returned to same position Success
Avg.
Baselines
Diffusion policy [ 43 ]
–
–
–
0
–
–
–
–
–
0
–
–
0
0.0
Diffusion policy [ 43 ] + history
15
10
0
0
100
25
0
0
25
0
15
–
0
0.0
X-VLA [ 42 ]
–
–
–
0
–
–
–
–
–
0
–
–
0
0.0
X-VLA [ 42 ] + history
0
10
85
10
100
0
35
5
40
5
75
65
10
8.3
TABLE IV: Real-robot results, step by step. Each task moves from observation-driven stages to a later decision that depends on retained history (outlined in Fig. 3 ). Breaking results down this way, instead of reporting only final success, shows exactly where a policy fails. For Plant , the three scoop columns are mutually exclusive: a rollout is counted under the number of scoops it performed before it either moved on to the flower or stayed in the scooping loop until timeout. For Pot Timer , “No wait” and “Wrong wait” are rollouts that left the waiting state early or outside the 29-31 s window. For Sponge , “Returned to wrong position” and “Returned to same position” partition the rollouts that placed the sponge, except rollouts that stalled at the plate. Success is the final column of each task. All values are percentages of 20 rollouts with randomised start positions. Avg. is the mean over the success rate of three tasks.
Plant
Pot timer
Sponge
Integration
1 scoop
2 scoops
3 or more
Flower planted Success
Pot on stove
No wait
Wrong wait
≈ 30 s wait
Pepper in the pot
Success
Pick sponge, put on plate
Returned to wrong position
Returned to same position
Avg.
ReCAT, cross-attention
0
100
0
85
100
0
0
100
95
95
20
0
20
66.7
Integration variants
Late fusion
0
0
0
0
90
0
0
90
90
90
70 †
40
25
38.3
AdaLN
35
15
0
5
100
0
0
100
90
90
25
5
0
31.7
Scale
0
25
0
10
80
0
0
80
75
75
40
40
0
28.3
TABLE V: Step-wise real-robot success across six memory-to-action conditioning mechanisms. Columns as in Table IV .
Setting
Sponge
Plant
Pot timer
Avg.
Ours, hybrid encoder
20
85
95
66.7
Ours, full-transformer encoder
10
0
65
25.0
Ours, pure-SSM encoder
–
25
–
–
Ours, w/o multimodal observation encoder tokens
0
0
0
0.0
Ours, without memory module
0
0
0
0.0
TABLE VI: Real-robot success rates for observation-encoder and component ablations.
Metric
ReCAT
DP-H
X-VLA-H
GMP
Total / trainable parameters
374M / 142M
267.9M / 267.9M
881.9M / 881.9M
282.2M / 188.7M
Peak inference VRAM
1.44GB
0.65GB
2.55GB
1.09GB
Inference latency/frequency
0.059s/16.9hz
0.11s/9hz
0.288s/3.47hz
0.093s/10.75hz
Smoothness metric (SPARC [ 41 ] )
-5.16
-6.19
-6.76
-13.06
TABLE VII: Computational cost on the same hardware and inference setting at. Latency is per policy inference; the rate is its reciprocal. SPARC: higher is smoother.
Memory-dependent robotic manipulation requires policies to use information that is no longer available in the current observation. Retaining history alone is insufficient: memory must preserve information that supports future actions. One challenge is whether a memory-free foundation model can learn to retain and use historical information from action demonstrations alone, without external memory support. We introduce T2Mem, a framework that develops this capability within a pretrained vision-language-action policy, without external reasoning models or memory-specific annotations. T2Mem uses test-time training to encode observation history into compact fast weights through online self-supervised updates, avoiding repeated processing of the full history. An observation-grounded interface extracts vision-language information for memory formation and supplies retrieved context to the action expert. Action supervision shapes what the memory learns to retain and use, while alternating memory-policy learning gives each component a fixed counterpart during optimization. Across 16 RoboMME tasks, T2Mem improves average success from 17.93% to 56.83% over the memory-free base policy and outperforms the recurrent-memory methods reported in the benchmark, while controlled profiling indicates at least 3x inference speedup over explicit methods. Project website: https://yzliu84.github.io/T2MEM-project/
Yize Liu, Huang Huang, Yining Hong +4
Stanford University · NVIDIA · University of Michigan, Ann Arbor
Vision-Language-Action models provide a strong foundation for general-purpose robot control, yet a vast majority of policies do not preserve and leverage episode-level information beyond the current observation. This limitation is consequential in history-dependent manipulation tasks that depend on information available only in past observations. Retaining past observations in context can aid in recovering this information, but at the significant cost of ever-growing, bloated context and inference latency. We thus introduce MemBodied, a fixed-size episodic memory with two complementary components: an associative state that records interactions across policy calls and an episode anchor that preserves a compact representation of the initial scene as a reference. At each policy call, the model conditions action generation on the current input and the memory components, rather than directly using past observations. Across five evaluated RMBench tasks requiring memory, MemBodied achieves 7.81× the mean success rate of a stateless policy and 2.98× of vanilla recurrent memory, while outperforming the strongest memory-augmented baseline by 1.3× with 10× fewer added parameters. On the fully observable LIBERO-Long suite, it reached 90.6%, a 5.4% improvement over the stateless π0 policy. These findings support MemBodied as a practical alternative to expanding the policy context for history-dependent manipulation.
Tej Deep Pala, Navonil Majumder, Bryce Goh +4
Nanyang Technological University · Griffin Labs · École Centrale de Lyon
Video-world-model policies learn action-relevant representations by predicting future observations. However, they condition on only a short observation window, which renders long-horizon manipulation non-Markovian when the correct action depends on earlier events that are no longer visible. We present MemoryVAM, an episodic memory mechanism for video-world-model policies. We employ a Recap-Cue (RC) module, in which a Perceiver-based Recap Compressor maps per-frame CLIP embeddings into compact memory tokens, and a lightweight Cue Gate estimates task completion from memory and language. These tokens are injected into both the video backbone and the action decoder, aligning policy imagination with episode progress and conditioning actions on history. Our model trains the memory module with video prediction, a delta-reconstruction auxiliary loss, and episode-boundary supervision, requiring no per-frame progress labels. The same mechanism applies to UNet and Diffusion Transformer (DiT) backbones by changing only the cross-attention injection interface. On LIBERO-Mem, our model improves average success from 5% to 42.5%. On real robots, it achieves 78.3% success on counting tasks, 80.0% on spatial recall, and 75.0% on sequential tracking. Project page: https://MemoryVAM.github.io/
Yuxin Jiang, Chang Yu, Yunuo Chen +4
University of California, Los Angeles · University of California, San Diego · University of Utah +1