cs.ROSep 28, 2026

ReCAT: Remember, Count, and Time: Structured Recurrent Memory for Robot Manipulation

Authors: Pankhuri Vanjani, Mostafa Hatab, Can Mizrakli, Vaisakh Shaj, Zhuoyue Li, Moritz Reuss, Rudolf Lioutikov

Organizations: Intuitive Robots Lab, Karlsruhe Institute of Technology (KIT), Germany · University of Edinburgh, UK · NVIDIA · Robotics Institute Germany (RIG)

Abstract

Memory-dependent manipulation requires robots to make decisions using information that is no longer available to their current sensors, such as recalling an earlier visual cue, tracking task progress, counting repeated events, or estimating elapsed time. We present ReCAT, a language-conditioned policy with structured recurrent memory. An instruction-conditioned encoder forms features from the current observation. A recurrent memory integrates the observation stream through Mamba-2 layers and one causal attention layer. A flow-matching Transformer decoder reads the current and the historical representation through separate cross-attention in every block. ReCAT reaches 95.3% average success on LIBERO and 62.4% on RMBench, with the best or tied-best result on six of nine tasks. On three real-robot tasks probing spatial recall, event counting, and interval timing, the best ReCAT variant reaches 66.7% average success, against 8.3% for the strongest short-history baseline. Controlled comparisons within ReCAT show that the observation encoder and every-block memory conditioning are needed for this performance. They also show that update rules developed for efficient sequence modeling behave differently as robot memory: additive updates have the highest observed success on counting and timing, and delta-rule updates on spatial recall. Project website is at https://intuitive-robots.github.io/ReCAT

Figures & tables

Explore similar work

CardsList
  1. T2^2Mem: Learning Test-Time Memory for Robotics

    Sep 29, 2026Yize Liu, Huang Huang, Yining Hong +4Robotic ManipulationTest Time

  2. MemBodied: Recurrent Associative Memory for Vision-Language-Action Models

    Sep 23, 2026Tej Deep Pala, Navonil Majumder, Bryce Goh +4Generalizable Vision-Language-Action PoliciesEmbodied

  3. MemoryVAM: Integrating Memory into Video Action Model for Robot Manipulation

    Jun 13, 2026Yuxin Jiang, Chang Yu, Yunuo Chen +4Efficient World-Action ModelRobotic Manipulation