Visual-memory systems commonly retain or compress past observations. Robot control additionally requires interaction-derived state that no individual frame may explicitly represent, such as persistent identity relations, accumulated progress, or ordered procedures. We introduce Simple Agentic Robot Memory (SimpleARM), a training-free memory layer for frozen generalist robot policies. From the task instruction, SimpleARM specifies what to monitor; frozen perceptual tools maintain compact typed state online; structured access retrieves that state only when a proposed subgoal depends on history; and current-view grounding resolves recalled entities before execution. We evaluate SimpleARM on RoboMME, a benchmark of memory-dependent robot manipulation tasks that require history information no longer available in the current observation. Across all 16 tasks and three policy seeds, SimpleARM achieves 67.17% mean success, compared with 44.51% for the strongest non-oracle baseline. Matched ablations show mechanism specificity: removing relation, reference, progress, or route state produces large losses where the affected state is retrieved for control, while largely sparing other tasks. These results support a state-based view of robot memory: effective memory for control is not simply retained visual history, but compact task-relevant state derived from the interaction history.
Figures & tables
Figure 1: SimpleARM overview. Given the instruction, SimpleARM first specifies what historical state may be needed, then maintains that state online with frozen perception and event tools, and retrieves it only when the current subgoal depends on interaction history. Retrieved memory resolves historical identity, progress, or procedure; current perception provides any required spatial grounding, and the resulting memory-conditioned grounded subgoal is executed by the frozen VLA.
success (%)
suite
task
required state
no mem.
MemER
FrameSamp
Recent32
ours
oracle
Counting
BinFill
progress
52.0
56.7
39.6
34.0
60.0
85.8
PickXtimes
progress
92.7
79.3
87.3
83.3
91.3
100.0
SwingXtimes
progress
7.3
59.3
92.0
80.7
76.7
100.0
StopCube
progress
0.0
0.0
42.0
41.3
0.0
49.7
suite avg.
ΔFS=−8.22
38.00
48.83
65.22
59.83
57.00
83.86
Table 1: Success rate (%) on RoboMME. All columns use three policy seeds and 50 episodes per task per seed. Published baseline and oracle results are from RoboMME ( Dai et al., 2026 ) , which also evaluates MemER ( Sridhar et al., 2026 ) ; Recent32+ModuL retains the latest 32 native frames (Sec. 4.1 ). ΔFS is ours minus FrameSamp. State labels are not mutually exclusive. Bold and underlining mark the best and second-best non-oracle results, respectively; ties share the same formatting. The ground-truth-subgoal oracle is shown in gray.
Figure 2: SimpleARM’s paired gains are concentrated in Permanence and Reference. (a) Success rates of SimpleARM and the released FrameSamp-Modul checkpoint on matched episode identities under the same seed (200 episodes per suite; 800 overall). Δ denotes the paired success-rate difference (SimpleARM minus FrameSamp), in percentage points. (b) Counts of episodes solved by exactly one system: FrameSamp alone (left) or SimpleARM alone (right). The bottom row in each panel pools all four suites.
Figure 3: Matched ablations localize to tasks that retrieve the state. (a) Cells show the matched change in success after removing one component (ablated minus full; 50 paired episodes per task). Solid borders mark tasks that retrieve the maintained state for control; dashed borders mark tasks where the mechanism maintains state in at least 20% of episodes but it is never retrieved. The identity-tools row uses a separate 450-episode evaluation over nine tasks; unevaluated cells are marked n/a . (b) Aggregate effects by retrieval status. Losses are large when state reaches control and small when it is maintained but never retrieved; diamonds denote access comparisons with no record to retrieve.
Figure 4: Representative failure cases. (a) SwingXtimes success by count accuracy. (b) PickHighlight success as initialized bindings increase. (c) Final carrier status for ButtonUnmask (BU) and ButtonUnmaskSwap (BUS). A carrier is a cover track bound to a remembered cube: “all tracks accepted” means every carrier ends internally verified or reacquired; “mixed” combines accepted and uncertain carriers; and “unresolved” means all carriers remain uncertain. These labels are internal states, not ground-truth identity checks. All panels pool three policy seeds and are diagnostic; bars in (a) and (b) are 95% Wilson intervals on the plotted rate.
Appendix figures & tables24 assets
Supplementary material from the paper’s appendix.
Appendix
state type
persistent field
verified update
policy-facing read
entity / relation
identity, role, container or current relation
detector proposal plus appearance/geometric consistency
historical referent; current geometry is re-grounded
event / progress
event count, repetition index or event-relative stage
task-observable event predicate
count, ordinal or completion field
trajectory / procedure
ordered route elements and progress index
parsed demonstration segment or verified route transition
next waypoint or ordered procedural element
Appendix
Table 2: Typed memory schemas in the RoboMME instantiation. Language selects fields to monitor; frozen tools supply and verify their values.
task
Recent4
Recent16
Recent32
Counting
BinFill
31.3
36.7
34.0
PickXtimes
74.0
80.0
83.3
SwingXtimes
83.3
78.0
80.7
StopCube
17.3
40.0
41.3
Permanence
Appendix
Table 3: Recent-frame baselines over three policy seeds. RoboMME’s FrameSamp+ModuL architecture trained and evaluated with the latest 4, 16, or 32 native frames. Each variant is trained once (seed 42) and its final checkpoint is evaluated with policy seeds 7, 11, and 23, 16 tasks × 50 episodes per seed. Task cells are three-seed mean success (%); the lower block gives each seed’s successes out of 800 and the mean and sample standard deviation across seeds. The Recent32 column is the one reported in Table 1 .
task
seed 7
seed 11
seed 23
mean
sd
BinFill
72
54
54
60.00
10.39
PickXtimes
88
92
94
91.33
3.06
SwingXtimes
78
76
76
76.67
1.15
StopCube
0
0
0
0.00
0.00
VideoUnmask
100
94
94
96.00
3.46
ButtonUnmask
76
90
80
82.00
7.21
Appendix
Table 4: Per-seed success (%) for the final method. Each seed contains 50 episodes for each of the sixteen tasks; the last column reports the sample standard deviation across seeds.
suite
task
history-dependent state
Counting
BinFill
count/progress
PickXtimes
count/progress
SwingXtimes
count/progress
StopCube ∗
count/progress
Permanence
VideoUnmask
entity–location relation
ButtonUnmask
entity–location relation
Appendix
Table 5: Task-state annotation derived from the oracle subgoal templates. Labels may overlap. ∗ marks state expressed only by the instruction, not by the oracle subgoal text.
task
ours
FS
Δ
only
task
ours
FS
Δ
only
Counting
Reference
BinFill
36/50
22/50
+28
17/3
PickHighlight
34/50
9/50
+50
26/1
PickXtimes
44/50
48/50
−8
1/5
VideoRepick
37/50
17/50
+40
22/2
SwingXtimes
39/50
44/50
−10
4/9
VideoPlaceButton
41/50
28/50
+26
15/2
StopCube
0/50
25/50
−50
0/25
VideoPlaceOrder
44/50
22/50
+44
23/1
suite
119/200
139/200
−10.0
22/42
suite
156/200
76/200
+40.0
86/6
Appendix
Table 6: Episode-matched comparison with the released FrameSamp-Modul checkpoint. The runs share task, scene, demonstration, and episode identity but were executed on different hosts. Tasks are grouped into the four RoboMME suites; both halves of the table have the same columns. FS abbreviates FrameSamp-Modul, Δ is ours minus FrameSamp in percentage points, and only counts the episodes solved by exactly one system, as ours-only / FrameSamp-only.
ablation
control
ablated
Δ
better/worse
Memory components
agent-decided access
67.38
55.50
−11.88
28/123
no route state
67.75
56.25
−11.50
24/116
no demo references
67.25
58.12
−9.12
43/116
no live identity track
67.50
61.62
−5.88
42/89
no demo notes
67.75
62.38
−5.38
38/81
Appendix
Table 7: All final-method ablations on the common 800-episode seed-11 pool. Each row compares the ablation with its episode-matched full-method control, ordered by the size of the effect. better/worse counts the matched pairs whose outcome changes in each direction, so their sum is the number of episodes the ablation altered at all. Row order matches Figure 5 .
Category
Ablation
Tasks requiring component
Δ required
Δ remaining
Maintained state
Route-state removal
PatternLock, RouteStick
−68.0
−3.4
Progress-count removal
SwingXtimes
−62.0
−1.2
Demonstration-reference removal
VideoRepick, VideoPlaceButton, VideoPlaceOrder
−44.7
−0.9
Live-binding removal
ButtonUnmask, ButtonUnmaskSwap
−53.0
+0.9
Demonstration-binding removal
VideoUnmask, VideoUnmaskSwap
−28.0
−2.1
State update
Demonstration relation update
VideoUnmaskSwap
−50.0
+0.3
Appendix
Table 8: Task specificity of memory-component ablations. Each ablation is evaluated on all 16 tasks with 50 matched episodes per task. Task sets are specified a priori from the state variable or access mechanism required by the final method. Δ denotes the change in success rate relative to the full method, in percentage points.
Figure 5: All ten final-method ablations on 800 episode-matched comparisons each. Points are ablated-minus-full success and lines are paired-bootstrap 95% intervals. The lower three rows are robustness controls rather than typed memory mechanisms.
BF
PX
SX
SC
VU
BU
VUS
BUS
PH
VR
VPB
VPO
MC
IP
PL
RS
no route
+4
0
−12
0
+2
−12
+2
−16
0
−8
0
−2
0
−6
−90
−46
no flash count
+10
−6
−62
0
+4
−2
0
−6
−2
−10
0
0
0
−6
−2
+2
no demo references
+6
−6
0
0
0
0
0
0
+2
−48
−34
−52
−2
−6
−4
−2
no live track
0
+6
0
0
+2
−64
+2
−42
+2
0
0
−2
+2
−4
−2
+6
no demo notes
+2
+2
−4
0
−8
−4
−48
−10
+4
−8
0
0
−6
−6
0
0
no demo following
+2
+2
−2
0
−6
+2
−50
+2
−4
0
0
−2
+10
−4
0
+4
Appendix
Table 9: Complete ablation–task matrix. Every cell contains 50 matched episodes at seed 11; entries are ablated minus full success (percentage points). Bold cells are predicted active before observing outcomes. Task abbreviations follow their column order in Table 1 .
Table 10: Mechanism activity does not imply contribution to control. Activity is the fraction of final-run episodes in which the mechanism is exercised; sensitive tasks are those affected by the corresponding ablation.
Task
Binding retrieved
Full
Ablated
Δ
ButtonUnmask
yes
42/50
9/50
−66
ButtonUnmaskSwap
yes
34/50
10/50
−48
VideoUnmaskSwap
yes
37/50
14/50
−46
VideoRepick
yes
34/50
16/50
−36
VideoUnmask
no
48/50
48/50
0
SwingXtimes
no
34/50
33/50
−2
Appendix
Table 11: Matched identity-tool ablation. The frozen feature anchor and optical-flow continuity check are removed while the color verifier is retained. The comparison covers nine tasks with 50 matched episodes each. “Retrieved” means that the maintained binding is read at decision time and conditions the action; the pooled “maintained, not retrieved” row is the open circle of Figure 3 b (the four tasks on which the tools run in at least 20% of episodes), and “all other tasks” adds VideoPlaceOrder, where they run in 8%. SAM is absent from both arms, so this comparison should not be interpreted as removing every identity tool.
Comparison
Tasks
Matched episodes
Δ
Instruction rewording
8
416
−17.8
with environment-provided evidence source
8
416
−7.7
Detector substitution
8
165
−5.5
Keyword rather than agent-specified write policy
11
422
−4.5
Write gate removed
11
548
−5.1
Structural access replaced by loose lexical access
6
222
−5.9
Appendix
Table 12: Supplementary robustness comparisons. Each row is matched by task and episode to a configuration-matched control. These comparisons use a distinct evaluation configuration and therefore characterize sensitivity rather than final-method component necessity. Δ is variant minus control in success percentage points.
scope
update condition
control → no flow
Δ
better/worse
ButtonUnmaskSwap
bound cover moves after occlusion
34/50→22/50
−24
2/14
other four tasks
no live post-cover update of this type
161/200→161/200
0
12/12
pooled
five predefined tasks
195/250→183/250
−4.8
14/26
Appendix
Table 13: Matched seed-7 no-flow control on five identity/update tasks. Both arms ran on a host where SAM was unavailable, unlike the final method; therefore this does not estimate flow’s marginal effect in the shipped tool stack. Δ is no-flow minus control, and better/worse counts the matched episodes whose outcome changes in each direction.
task
required policy variable
represented?
evidence
SwingXtimes
accumulated swing count
yes
no count: −62 points
ButtonUnmask
live cube–container binding
no binding: −64 points
VideoPlaceOrder
demonstrated ordinal target
no reference: −52 points
PatternLock
ordered route
no route: −90 points
RouteStick
ordered route and turn sense
53.3%, near 55.6% oracle ceiling
StopCube
timed reach count of moving cube
no
0/150; no state formed
Appendix
Table 14: Representational coverage of the current memory schema. Matched removals quantify the contribution of represented variables; uncovered tasks require a control variable absent from the schema. “Represented” does not imply that the complete controller succeeds.
task
failures
furthest recorded attainment (episodes)
task-level interpretation
BinFill
60
state formed, no recorded access: 60
count state maintained
PickXtimes
13
state formed, no recorded access: 13
count state maintained
SwingXtimes
35
state formed; ordinal access analyzed separately
progress state maintained
StopCube
150
no usable typed state: 150
required timed state absent
VideoUnmask
6
state returned: 5; formed without access: 1
demonstration relation
ButtonUnmask
27
state returned: 27
live entity–carrier binding
Appendix
Table 15: Memory-state attainment among the 788 failed episodes. Each count records the furthest stage observed at any point in the episode, not the stage that caused failure. Specialized reference access in VideoPlaceButton and VideoPlaceOrder is not represented by the same access statistic and is therefore reported separately.
PatternLock
RouteStick
route length
n
success
n
success
1
24
100.0%
—
—
2
27
100.0%
42
76.2%
3
54
100.0%
36
58.3%
4
24
91.7%
27
40.7%
5
3
33.3%
27
25.9%
Appendix
Table 16: Final-method success by parsed route length. These are observational difficulty slices, not controlled ablations; small non-monotonic cells, especially length seven, should not be overread.
Figure 6: A complete ButtonUnmaskSwap memory lifecycle (seed 7, episode 47). The memory writes cube identities while they are visible, converts them to cube–cover bindings after occlusion, and updates the bindings as the covers move. At both history-dependent reads, the frozen composer proposes the wrong container; the memory retrieves the corresponding binding, re-grounds it to current geometry, and sends the corrected coordinate to the frozen VLA. The bottom lanes show the composer proposal, memory operations, and coordinate used for control.
Figure 7: Live entity binding, cover following, structured access, and current-view re-grounding on ButtonUnmaskSwap.
Figure 8: Demonstration-conditioned notes and cover following on VideoUnmaskSwap.
Figure 9: Verified flash counting and ordinal access on SwingXtimes.
Figure 10: Demonstration event reading, entity anchoring, and re-grounding on VideoRepick.
Figure 11: Ordinal placement state extracted from the demonstration on VideoPlaceOrder.
Figure 12: An ordered route parsed from demonstration cues and retrieved segment by segment on PatternLock.
Figure 13: A transient highlight maintained as a persistent entity–mark relation on PickHighlight.