Visual-memory systems commonly retain or compress past observations. Robot control additionally requires interaction-derived state that no individual frame may explicitly represent, such as persistent identity relations, accumulated progress, or ordered procedures. We introduce Simple Agentic Robot Memory (SimpleARM), a training-free memory layer for frozen generalist robot policies. From the task instruction, SimpleARM specifies what to monitor; frozen perceptual tools maintain compact typed state online; structured access retrieves that state only when a proposed subgoal depends on history; and current-view grounding resolves recalled entities before execution. We evaluate SimpleARM on RoboMME, a benchmark of memory-dependent robot manipulation tasks that require history information no longer available in the current observation. Across all 16 tasks and three policy seeds, SimpleARM achieves 67.17% mean success, compared with 44.51% for the strongest non-oracle baseline. Matched ablations show mechanism specificity: removing relation, reference, progress, or route state produces large losses where the affected state is retrieved for control, while largely sparing other tasks. These results support a state-based view of robot memory: effective memory for control is not simply retained visual history, but compact task-relevant state derived from the interaction history.
Figures & tables
Figure 1: SimpleARM overview. Given the instruction, SimpleARM first specifies what historical state may be needed, then maintains that state online with frozen perception and event tools, and retrieves it only when the current subgoal depends on interaction history. Retrieved memory resolves historical identity, progress, or procedure; current perception provides any required spatial grounding, and the resulting memory-conditioned grounded subgoal is executed by the frozen VLA.
success (%)
suite
task
required state
no mem.
MemER
FrameSamp
Recent32
ours
oracle
Counting
BinFill
progress
52.0
56.7
39.6
34.0
60.0
85.8
PickXtimes
progress
92.7
79.3
87.3
83.3
91.3
100.0
SwingXtimes
progress
7.3
59.3
92.0
80.7
76.7
100.0
StopCube
progress
0.0
0.0
42.0
41.3
0.0
49.7
suite avg.
ΔFS=−8.22
38.00
48.83
65.22
59.83
57.00
83.86
Table 1: Success rate (%) on RoboMME. All columns use three policy seeds and 50 episodes per task per seed. Published baseline and oracle results are from RoboMME ( Dai et al., 2026 ) , which also evaluates MemER ( Sridhar et al., 2026 ) ; Recent32+ModuL retains the latest 32 native frames (Sec. 4.1 ). ΔFS is ours minus FrameSamp. State labels are not mutually exclusive. Bold and underlining mark the best and second-best non-oracle results, respectively; ties share the same formatting. The ground-truth-subgoal oracle is shown in gray.
Figure 2: SimpleARM’s paired gains are concentrated in Permanence and Reference. (a) Success rates of SimpleARM and the released FrameSamp-Modul checkpoint on matched episode identities under the same seed (200 episodes per suite; 800 overall). Δ denotes the paired success-rate difference (SimpleARM minus FrameSamp), in percentage points. (b) Counts of episodes solved by exactly one system: FrameSamp alone (left) or SimpleARM alone (right). The bottom row in each panel pools all four suites.
Figure 3: Matched ablations localize to tasks that retrieve the state. (a) Cells show the matched change in success after removing one component (ablated minus full; 50 paired episodes per task). Solid borders mark tasks that retrieve the maintained state for control; dashed borders mark tasks where the mechanism maintains state in at least 20% of episodes but it is never retrieved. The identity-tools row uses a separate 450-episode evaluation over nine tasks; unevaluated cells are marked n/a . (b) Aggregate effects by retrieval status. Losses are large when state reaches control and small when it is maintained but never retrieved; diamonds denote access comparisons with no record to retrieve.
Figure 4: Representative failure cases. (a) SwingXtimes success by count accuracy. (b) PickHighlight success as initialized bindings increase. (c) Final carrier status for ButtonUnmask (BU) and ButtonUnmaskSwap (BUS). A carrier is a cover track bound to a remembered cube: “all tracks accepted” means every carrier ends internally verified or reacquired; “mixed” combines accepted and uncertain carriers; and “unresolved” means all carriers remain uncertain. These labels are internal states, not ground-truth identity checks. All panels pool three policy seeds and are diagnostic; bars in (a) and (b) are 95% Wilson intervals on the plotted rate.
Appendix figures & tables24 assets
Supplementary material from the paper’s appendix.
Appendix
state type
persistent field
verified update
policy-facing read
entity / relation
identity, role, container or current relation
detector proposal plus appearance/geometric consistency
historical referent; current geometry is re-grounded
event / progress
event count, repetition index or event-relative stage
task-observable event predicate
count, ordinal or completion field
trajectory / procedure
ordered route elements and progress index
parsed demonstration segment or verified route transition
next waypoint or ordered procedural element
Appendix
Table 2: Typed memory schemas in the RoboMME instantiation. Language selects fields to monitor; frozen tools supply and verify their values.
task
Recent4
Recent16
Recent32
Counting
BinFill
31.3
36.7
34.0
PickXtimes
74.0
80.0
83.3
SwingXtimes
83.3
78.0
80.7
StopCube
17.3
40.0
41.3
Permanence
Appendix
Table 3: Recent-frame baselines over three policy seeds. RoboMME’s FrameSamp+ModuL architecture trained and evaluated with the latest 4, 16, or 32 native frames. Each variant is trained once (seed 42) and its final checkpoint is evaluated with policy seeds 7, 11, and 23, 16 tasks × 50 episodes per seed. Task cells are three-seed mean success (%); the lower block gives each seed’s successes out of 800 and the mean and sample standard deviation across seeds. The Recent32 column is the one reported in Table 1 .
task
seed 7
seed 11
seed 23
mean
sd
BinFill
72
54
54
60.00
10.39
PickXtimes
88
92
94
91.33
3.06
SwingXtimes
78
76
76
76.67
1.15
StopCube
0
0
0
0.00
0.00
VideoUnmask
100
94
94
96.00
3.46
ButtonUnmask
76
90
80
82.00
7.21
Appendix
Table 4: Per-seed success (%) for the final method. Each seed contains 50 episodes for each of the sixteen tasks; the last column reports the sample standard deviation across seeds.
suite
task
history-dependent state
Counting
BinFill
count/progress
PickXtimes
count/progress
SwingXtimes
count/progress
StopCube ∗
count/progress
Permanence
VideoUnmask
entity–location relation
ButtonUnmask
entity–location relation
Appendix
Table 5: Task-state annotation derived from the oracle subgoal templates. Labels may overlap. ∗ marks state expressed only by the instruction, not by the oracle subgoal text.
task
ours
FS
Δ
only
task
ours
FS
Δ
only
Counting
Reference
BinFill
36/50
22/50
+28
17/3
PickHighlight
34/50
9/50
+50
26/1
PickXtimes
44/50
48/50
−8
1/5
VideoRepick
37/50
17/50
+40
22/2
SwingXtimes
39/50
44/50
−10
4/9
VideoPlaceButton
41/50
28/50
+26
15/2
StopCube
0/50
25/50
−50
0/25
VideoPlaceOrder
44/50
22/50
+44
23/1
suite
119/200
139/200
−10.0
22/42
suite
156/200
76/200
+40.0
86/6
Appendix
Table 6: Episode-matched comparison with the released FrameSamp-Modul checkpoint. The runs share task, scene, demonstration, and episode identity but were executed on different hosts. Tasks are grouped into the four RoboMME suites; both halves of the table have the same columns. FS abbreviates FrameSamp-Modul, Δ is ours minus FrameSamp in percentage points, and only counts the episodes solved by exactly one system, as ours-only / FrameSamp-only.
ablation
control
ablated
Δ
better/worse
Memory components
agent-decided access
67.38
55.50
−11.88
28/123
no route state
67.75
56.25
−11.50
24/116
no demo references
67.25
58.12
−9.12
43/116
no live identity track
67.50
61.62
−5.88
42/89
no demo notes
67.75
62.38
−5.38
38/81
Appendix
Table 7: All final-method ablations on the common 800-episode seed-11 pool. Each row compares the ablation with its episode-matched full-method control, ordered by the size of the effect. better/worse counts the matched pairs whose outcome changes in each direction, so their sum is the number of episodes the ablation altered at all. Row order matches Figure 5 .
Category
Ablation
Tasks requiring component
Δ required
Δ remaining
Maintained state
Route-state removal
PatternLock, RouteStick
−68.0
−3.4
Progress-count removal
SwingXtimes
−62.0
−1.2
Demonstration-reference removal
VideoRepick, VideoPlaceButton, VideoPlaceOrder
−44.7
−0.9
Live-binding removal
ButtonUnmask, ButtonUnmaskSwap
−53.0
+0.9
Demonstration-binding removal
VideoUnmask, VideoUnmaskSwap
−28.0
−2.1
State update
Demonstration relation update
VideoUnmaskSwap
−50.0
+0.3
Appendix
Table 8: Task specificity of memory-component ablations. Each ablation is evaluated on all 16 tasks with 50 matched episodes per task. Task sets are specified a priori from the state variable or access mechanism required by the final method. Δ denotes the change in success rate relative to the full method, in percentage points.
Figure 5: All ten final-method ablations on 800 episode-matched comparisons each. Points are ablated-minus-full success and lines are paired-bootstrap 95% intervals. The lower three rows are robustness controls rather than typed memory mechanisms.
BF
PX
SX
SC
VU
BU
VUS
BUS
PH
VR
VPB
VPO
MC
IP
PL
RS
no route
+4
0
−12
0
+2
−12
+2
−16
0
−8
0
−2
0
−6
−90
−46
no flash count
+10
−6
−62
0
+4
−2
0
−6
−2
−10
0
0
0
−6
−2
+2
no demo references
+6
−6
0
0
0
0
0
0
+2
−48
−34
−52
−2
−6
−4
−2
no live track
0
+6
0
0
+2
−64
+2
−42
+2
0
0
−2
+2
−4
−2
+6
no demo notes
+2
+2
−4
0
−8
−4
−48
−10
+4
−8
0
0
−6
−6
0
0
no demo following
+2
+2
−2
0
−6
+2
−50
+2
−4
0
0
−2
+10
−4
0
+4
Appendix
Table 9: Complete ablation–task matrix. Every cell contains 50 matched episodes at seed 11; entries are ablated minus full success (percentage points). Bold cells are predicted active before observing outcomes. Task abbreviations follow their column order in Table 1 .
Table 10: Mechanism activity does not imply contribution to control. Activity is the fraction of final-run episodes in which the mechanism is exercised; sensitive tasks are those affected by the corresponding ablation.
Task
Binding retrieved
Full
Ablated
Δ
ButtonUnmask
yes
42/50
9/50
−66
ButtonUnmaskSwap
yes
34/50
10/50
−48
VideoUnmaskSwap
yes
37/50
14/50
−46
VideoRepick
yes
34/50
16/50
−36
VideoUnmask
no
48/50
48/50
0
SwingXtimes
no
34/50
33/50
−2
Appendix
Table 11: Matched identity-tool ablation. The frozen feature anchor and optical-flow continuity check are removed while the color verifier is retained. The comparison covers nine tasks with 50 matched episodes each. “Retrieved” means that the maintained binding is read at decision time and conditions the action; the pooled “maintained, not retrieved” row is the open circle of Figure 3 b (the four tasks on which the tools run in at least 20% of episodes), and “all other tasks” adds VideoPlaceOrder, where they run in 8%. SAM is absent from both arms, so this comparison should not be interpreted as removing every identity tool.
Comparison
Tasks
Matched episodes
Δ
Instruction rewording
8
416
−17.8
with environment-provided evidence source
8
416
−7.7
Detector substitution
8
165
−5.5
Keyword rather than agent-specified write policy
11
422
−4.5
Write gate removed
11
548
−5.1
Structural access replaced by loose lexical access
6
222
−5.9
Appendix
Table 12: Supplementary robustness comparisons. Each row is matched by task and episode to a configuration-matched control. These comparisons use a distinct evaluation configuration and therefore characterize sensitivity rather than final-method component necessity. Δ is variant minus control in success percentage points.
scope
update condition
control → no flow
Δ
better/worse
ButtonUnmaskSwap
bound cover moves after occlusion
34/50→22/50
−24
2/14
other four tasks
no live post-cover update of this type
161/200→161/200
0
12/12
pooled
five predefined tasks
195/250→183/250
−4.8
14/26
Appendix
Table 13: Matched seed-7 no-flow control on five identity/update tasks. Both arms ran on a host where SAM was unavailable, unlike the final method; therefore this does not estimate flow’s marginal effect in the shipped tool stack. Δ is no-flow minus control, and better/worse counts the matched episodes whose outcome changes in each direction.
task
required policy variable
represented?
evidence
SwingXtimes
accumulated swing count
yes
no count: −62 points
ButtonUnmask
live cube–container binding
no binding: −64 points
VideoPlaceOrder
demonstrated ordinal target
no reference: −52 points
PatternLock
ordered route
no route: −90 points
RouteStick
ordered route and turn sense
53.3%, near 55.6% oracle ceiling
StopCube
timed reach count of moving cube
no
0/150; no state formed
Appendix
Table 14: Representational coverage of the current memory schema. Matched removals quantify the contribution of represented variables; uncovered tasks require a control variable absent from the schema. “Represented” does not imply that the complete controller succeeds.
task
failures
furthest recorded attainment (episodes)
task-level interpretation
BinFill
60
state formed, no recorded access: 60
count state maintained
PickXtimes
13
state formed, no recorded access: 13
count state maintained
SwingXtimes
35
state formed; ordinal access analyzed separately
progress state maintained
StopCube
150
no usable typed state: 150
required timed state absent
VideoUnmask
6
state returned: 5; formed without access: 1
demonstration relation
ButtonUnmask
27
state returned: 27
live entity–carrier binding
Appendix
Table 15: Memory-state attainment among the 788 failed episodes. Each count records the furthest stage observed at any point in the episode, not the stage that caused failure. Specialized reference access in VideoPlaceButton and VideoPlaceOrder is not represented by the same access statistic and is therefore reported separately.
PatternLock
RouteStick
route length
n
success
n
success
1
24
100.0%
—
—
2
27
100.0%
42
76.2%
3
54
100.0%
36
58.3%
4
24
91.7%
27
40.7%
5
3
33.3%
27
25.9%
Appendix
Table 16: Final-method success by parsed route length. These are observational difficulty slices, not controlled ablations; small non-monotonic cells, especially length seven, should not be overread.
Figure 6: A complete ButtonUnmaskSwap memory lifecycle (seed 7, episode 47). The memory writes cube identities while they are visible, converts them to cube–cover bindings after occlusion, and updates the bindings as the covers move. At both history-dependent reads, the frozen composer proposes the wrong container; the memory retrieves the corresponding binding, re-grounds it to current geometry, and sends the corrected coordinate to the frozen VLA. The bottom lanes show the composer proposal, memory operations, and coordinate used for control.
Figure 7: Live entity binding, cover following, structured access, and current-view re-grounding on ButtonUnmaskSwap.
Figure 8: Demonstration-conditioned notes and cover following on VideoUnmaskSwap.
Figure 9: Verified flash counting and ordinal access on SwingXtimes.
Figure 10: Demonstration event reading, entity anchoring, and re-grounding on VideoRepick.
Figure 11: Ordinal placement state extracted from the demonstration on VideoPlaceOrder.
Figure 12: An ordered route parsed from demonstration cues and retrieved segment by segment on PatternLock.
Figure 13: A transient highlight maintained as a persistent entity–mark relation on PickHighlight.
A robot may lose sight of an object it must later retrieve, need to recall what a person demonstrated earlier, or track which steps of a task it has already completed. Current vision-language-action (VLA) policies often fail once the information needed for action disappears from the current observation, making memory critical for long-horizon robot behavior. Existing approaches typically provide longer histories or learn implicit memory from observation-action trajectories. But action supervision tells a policy how to act, not what to remember: it does not specify which past facts should persist or how they should change as new evidence arrives. We therefore separate maintaining an evidence-grounded account of the past from learning how to act on it. This insight motivates Explicit Concept Memory (ECoMEM), which represents task-relevant history with a reusable library of grounded concepts. An evidence-based Writer selects and updates these records, while a learned Reader turns them into memory tokens that directly condition the VLA. Across 16 RoboMME tasks, ECoMEM leads the evaluated robot policies on 15 tasks. On two new real-robot tasks, the same memory library either transfers directly or requires only one new concept, achieving 86.1% success versus 8.6% for a no-memory VLA. These results show that explicit concepts provide a reusable and extensible memory interface for robot control. Project website: https://ecomem.github.io/
Vision-Language-Action models provide a strong foundation for general-purpose robot control, yet a vast majority of policies do not preserve and leverage episode-level information beyond the current observation. This limitation is consequential in history-dependent manipulation tasks that depend on information available only in past observations. Retaining past observations in context can aid in recovering this information, but at the significant cost of ever-growing, bloated context and inference latency. We thus introduce MemBodied, a fixed-size episodic memory with two complementary components: an associative state that records interactions across policy calls and an episode anchor that preserves a compact representation of the initial scene as a reference. At each policy call, the model conditions action generation on the current input and the memory components, rather than directly using past observations. Across five evaluated RMBench tasks requiring memory, MemBodied achieves 7.81× the mean success rate of a stateless policy and 2.98× of vanilla recurrent memory, while outperforming the strongest memory-augmented baseline by 1.3× with 10× fewer added parameters. On the fully observable LIBERO-Long suite, it reached 90.6%, a 5.4% improvement over the stateless π0 policy. These findings support MemBodied as a practical alternative to expanding the policy context for history-dependent manipulation.
Tej Deep Pala, Navonil Majumder, Bryce Goh +4
Nanyang Technological University · Griffin Labs · École Centrale de Lyon
Memory-dependent robotic manipulation requires policies to use information that is no longer available in the current observation. Retaining history alone is insufficient: memory must preserve information that supports future actions. One challenge is whether a memory-free foundation model can learn to retain and use historical information from action demonstrations alone, without external memory support. We introduce T2Mem, a framework that develops this capability within a pretrained vision-language-action policy, without external reasoning models or memory-specific annotations. T2Mem uses test-time training to encode observation history into compact fast weights through online self-supervised updates, avoiding repeated processing of the full history. An observation-grounded interface extracts vision-language information for memory formation and supplies retrieved context to the action expert. Action supervision shapes what the memory learns to retain and use, while alternating memory-policy learning gives each component a fixed counterpart during optimization. Across 16 RoboMME tasks, T2Mem improves average success from 17.93% to 56.83% over the memory-free base policy and outperforms the recurrent-memory methods reported in the benchmark, while controlled profiling indicates at least 3x inference speedup over explicit methods. Project website: https://yzliu84.github.io/T2MEM-project/
Yize Liu, Huang Huang, Yining Hong +4
Stanford University · NVIDIA · University of Michigan, Ann Arbor