ReCAT: Remember, Count, and Time: Structured Recurrent Memory for Robot Manipulation
Authors: Pankhuri Vanjani, Mostafa Hatab, Can Mizrakli, Vaisakh Shaj, Zhuoyue Li, Moritz Reuss, Rudolf Lioutikov
Organizations: Intuitive Robots Lab, Karlsruhe Institute of Technology (KIT), Germany · University of Edinburgh, UK · NVIDIA · Robotics Institute Germany (RIG)
Memory-dependent manipulation requires robots to make decisions using information that is no longer available to their current sensors, such as recalling an earlier visual cue, tracking task progress, counting repeated events, or estimating elapsed time. We present ReCAT, a language-conditioned policy with structured recurrent memory. An instruction-conditioned encoder forms features from the current observation. A recurrent memory integrates the observation stream through Mamba-2 layers and one causal attention layer. A flow-matching Transformer decoder reads the current and the historical representation through separate cross-attention in every block. ReCAT reaches 95.3% average success on LIBERO and 62.4% on RMBench, with the best or tied-best result on six of nine tasks. On three real-robot tasks probing spatial recall, event counting, and interval timing, the best ReCAT variant reaches 66.7% average success, against 8.3% for the strongest short-history baseline. Controlled comparisons within ReCAT show that the observation encoder and every-block memory conditioning are needed for this performance. They also show that update rules developed for efficient sequence modeling behave differently as robot memory: additive updates have the highest observed success on counting and timing, and delta-rule updates on spatial recall. Project website is at https://intuitive-robots.github.io/ReCAT
Figures & tables
Fig. 1: Visually similar observations can require different actions depending on task history. ReCAT retains the completed scoop count in a compact recurrent memory, allowing the robot to scoop again after one scoop and put down the shovel after two.
Fig. 2: Overview of ReCAT. Frozen DynaFLIP and T5 encoders (LoRA-adapted) feed a language-conditioned visual compressor and a six-layer Mamba–attention frame encoder, which produces one feature vector zt per step. The same zt feeds the action decoder directly and is the input to the recurrent memory, which returns the readout mt . The flow-matching decoder attends to zt and mt through separate cross-attention in every block. The decoder backbone is shared across all experiments. The ablations vary the frame-encoder mixer, the recurrent update rule, memory depth and width, and how mt enters the decoder.
Model
State update
Update semantics
Mamba-2
diag(at)ht−1+vtkt⊤
Accumulate: storing a similar cue again adds another write to its trace, so repeated occurrences of an event build up in the state, the property counting requires.
Mamba-3
diag(ateiθt)ht−1+vtkt⊤
Accumulate + rotate: as above, with phase evolution that provides an additional signal for temporal progression or motion.
GDN-2
αt(I−βtktkt⊤)ht−1+βtvtkt⊤
Replace: storing a similar cue overwrites its previous content with the latest value, while other stored cues remain undisturbed. Repeated identical events therefore leave a single trace.
TABLE I: Update rules of the recurrent layers compared in Sec. IV . Each step writes a cue-content pair (kt,vt) derived from zt into the state ht . The rules differ in what the write does to information already stored. The descriptions give the intended semantics of each rule. Whether a trained policy uses them this way is an empirical question.
Fig. 3: Real-robot tasks, one per row, each isolating one memory demand: plant counts scoops across five stages, pot timer times an interval while the scene is static, and sponge recalls an unmarked start position that has left the camera view. Outlined frames mark the decision point where the correct action depends on an earlier event, not on the current observation.
Method
Spatial
Object
Goal
LIBERO-10
Avg.
LIBERO-Plus
Diffusion policy
78.3
92.5
68.3
50.5
72.4
–
X-VLA [ 42 ]
98.2
98.6
97.8
97.6
98.1
71.4
Ours
96.4
97.2
97.2
90.4
95.3
64.2
TABLE II: Success rates (%) on LIBERO.
Task
TMC
DP (90M)
ACT (80M)
\/∗pi0.5 (3B)
*X-VLA (0.9B)
MEM-0 (10B)
EventVLA (4B)
*MemoryVLA (7.3B)
*MemER (10B)
Ours (374M)
Observe and Pick Up
M(1)
1
1
9
9
4
21
2
7
4
Rearrange Blocks
M(1)
0
29
13
13
89
96
53
17
96
Put Back Block
M(1)
0
0
11
18
90
95
81
0
100
Swap Blocks
M(1)
11
2
24
16
67
96
76
14
100
Swap T
M(1)
20
2
15
3
14
87
9
7
64
Average
M(1)
6.4
6.8
14.4
11.8
52.8
79.0
44.2
9.0
72.8
TABLE III: Success rates (%) on RMBench. * Indicates models utilizing large scale pretraining.
Plant
Pot timer
Sponge
Method
1 scoop
2 scoops
3 or more
Flower planted Success
Pot on stove
No wait
Wrong wait
≈ 30 s wait
Pepper in the pot
Success
Pick sponge, put on plate
Returned to wrong position
Returned to same position Success
Avg.
Baselines
Diffusion policy [ 43 ]
–
–
–
0
–
–
–
–
–
0
–
–
0
0.0
Diffusion policy [ 43 ] + history
15
10
0
0
100
25
0
0
25
0
15
–
0
0.0
X-VLA [ 42 ]
–
–
–
0
–
–
–
–
–
0
–
–
0
0.0
X-VLA [ 42 ] + history
0
10
85
10
100
0
35
5
40
5
75
65
10
8.3
TABLE IV: Real-robot results, step by step. Each task moves from observation-driven stages to a later decision that depends on retained history (outlined in Fig. 3 ). Breaking results down this way, instead of reporting only final success, shows exactly where a policy fails. For Plant , the three scoop columns are mutually exclusive: a rollout is counted under the number of scoops it performed before it either moved on to the flower or stayed in the scooping loop until timeout. For Pot Timer , “No wait” and “Wrong wait” are rollouts that left the waiting state early or outside the 29-31 s window. For Sponge , “Returned to wrong position” and “Returned to same position” partition the rollouts that placed the sponge, except rollouts that stalled at the plate. Success is the final column of each task. All values are percentages of 20 rollouts with randomised start positions. Avg. is the mean over the success rate of three tasks.
Plant
Pot timer
Sponge
Integration
1 scoop
2 scoops
3 or more
Flower planted Success
Pot on stove
No wait
Wrong wait
≈ 30 s wait
Pepper in the pot
Success
Pick sponge, put on plate
Returned to wrong position
Returned to same position
Avg.
ReCAT, cross-attention
0
100
0
85
100
0
0
100
95
95
20
0
20
66.7
Integration variants
Late fusion
0
0
0
0
90
0
0
90
90
90
70 †
40
25
38.3
AdaLN
35
15
0
5
100
0
0
100
90
90
25
5
0
31.7
Scale
0
25
0
10
80
0
0
80
75
75
40
40
0
28.3
TABLE V: Step-wise real-robot success across six memory-to-action conditioning mechanisms. Columns as in Table IV .
Setting
Sponge
Plant
Pot timer
Avg.
Ours, hybrid encoder
20
85
95
66.7
Ours, full-transformer encoder
10
0
65
25.0
Ours, pure-SSM encoder
–
25
–
–
Ours, w/o multimodal observation encoder tokens
0
0
0
0.0
Ours, without memory module
0
0
0
0.0
TABLE VI: Real-robot success rates for observation-encoder and component ablations.
Metric
ReCAT
DP-H
X-VLA-H
GMP
Total / trainable parameters
374M / 142M
267.9M / 267.9M
881.9M / 881.9M
282.2M / 188.7M
Peak inference VRAM
1.44GB
0.65GB
2.55GB
1.09GB
Inference latency/frequency
0.059s/16.9hz
0.11s/9hz
0.288s/3.47hz
0.093s/10.75hz
Smoothness metric (SPARC [ 41 ] )
-5.16
-6.19
-6.76
-13.06
TABLE VII: Computational cost on the same hardware and inference setting at. Latency is per policy inference; the rate is its reciprocal. SPARC: higher is smoother.