Frozen vision-language-action (VLA) policies offer broad manipulation skills but execute open-loop action chunks without tracking task progress, so the agent cannot reliably decide whether to continue, retry, or terminate. External memory is a natural remedy, yet we find it can be harmful when attempted actions are recorded as completed progress: transient execution failures become persistent task-state errors, and such a memory can underperform no progress memory at all. We propose Achievement-Grounded Memory (AGM), a lightweight closed-loop framework for frozen VLA policies. AGM represents a task as a static subgoal sequence with a dynamic progress pointer and advances the pointer on physically verified achievement rather than on attempts. Proprioceptive gripper-load cues decide when to verify; coherent point tracking verifies grasps, and language-conditioned cross-view comparison, read by a single trained 2.43M-parameter verification head, verifies placements. The policy, tracker, and encoder remain frozen, the head is the only trained component, and deployment needs no auxiliary vision-language model. On the RoboMME Counting benchmark, AGM reaches 100.0% on PickXTimes and 84.0% on BinFill, surpassing the strongest memory-augmented baseline by 7.7 points on the four-task average, and the gains carry over to a physical robot, where AGM reaches 100.0% and 82.0%. These results suggest that reliable embodied memory depends more on disciplined state updates than on memory capacity.
Figures & tables
Figure 1: Overview of AGM. Top: the instruction expands into a typed subgoal sequence G , which with a progress pointer forms the task memory Mt=(G,pt) . The frozen VLA acts on gpt ; a gripper-load state machine detects interaction events, and only events compatible with the current subgoal type trigger verification, since an interaction alone is not evidence of achievement. The verdict updates the pointer by Eq. ( 13 ); the colored arrows depict that rule, not three outcomes reachable at one subgoal. Bottom: grasp achievement is verified by coherent upward point motion under a frozen tracker, placement by language-conditioned pre/post-release cross-view evidence under a frozen encoder. Snowflakes mark frozen modules; only the 2.43M verification head (flame) is trained.
Figure 2: Error correction in simulation. The instruction requires three blue and two green cubes in the bin. The second grasp fails: verification rejects it, the pointer is retained, and the policy retries successfully, whereas an attempt-based memory would have recorded the failure as progress and finished one cube short.
Success Rate (%)
Params.
Method
PickXTimes
BinFill
SwingXTimes
StopCube
Avg.
External
Trainable
Frozen π0.5
42.89
30.00
35.56
6.67
28.78
0
0
π0.5 w/ past actions
58.33
26.67
26.67
4.67
29.09
0
≈ 3.35B
SAM2Act+ ‡
76.00
40.00
25.33
0.00
35.33
–
n/r
SimpleSG + Qwen3-VL-4B
95.30
77.60
5.10
0.40
44.60
4B
LoRA ( n/r )
SimpleSG + Gemini-2.5-Pro
63.00
46.00
45.00
2.00
39.00
API
0
Table 1: Success rates (%) on the RoboMME Counting test split (50 episodes per task; AGM averaged over two runs). External : frozen parameters added beyond the 3.35B π0.5 backbone; Trainable : parameters updated per method. † : subgoals not anchored to gripper interactions, where AGM follows the frozen policy. ‡ : independent policy not built on π0.5 . API: proprietary model of undisclosed size; n/r : not reported. Best per column in bold .
PickXTimes
BinFill
Method
N=1
N=2
N=3
N=4
N=5
N=1
N=2
N=3
N=4
N=5
Sim
π0†
12.5
0.0
0.0
0.0
0.0
0.0
0.0
0.0
0.0
0.0
π0 +AGM
50.0
40.0
33.3
50.0
50.0
30.0
10.0
0.0
0.0
0.0
π0.5
100.0
30.0
16.7
0.0
0.0
90.0
20.0
18.8
12.5
0.0
π0.5 +AGM
100.0
100.0
100.0
100.0
100.0
100.0
95.0
84.4
62.5
66.7
Real
π0
20.0
0.0
0.0
0.0
0.0
0.0
0.0
0.0
0.0
0.0
Table 2: Success rates (%) by instructed repetition count N . Sim rows use the official test split (16/10/12/6/6 and 10/10/16/8/6 episodes per N for PickXTimes and BinFill; episode-weighted means match Table 1 ); π0.5 is our reproduction of the official checkpoint and π0 a subgoal-conditioned LoRA, marked † since full instructions are out-of-distribution for it. Real rows use 10 trials per count with policies and verification head retrained for the platform under an unchanged framework. Best per column in bold .
Figure 3: Real-robot rollouts. Filling the bin with three red cubes, placing a blue cube on the target three times, and filling the bin with three green cubes while a human replenishes cubes mid-task; the heart-shaped initial layout (top) and the mid-task replenishment (bottom) do not appear in the demonstrations. Overlays show the subgoal issued at each moment.
Grasp verification
Thresholds
PickXTimes
BinFill
CoTracker (full)
reversibility-aware
100.0
84.0
Proprioception only
reversibility-aware
98.0
78.0
CoTracker
flattened
98.0
74.0
Table 3: Ablations on grasp-verification modality and reversibility-aware thresholds (success rate %, official test split, n=50 per cell, same frozen π0.5 and head as Table 1 ; row 1 is the full method, BinFill as mean of two runs). Proprioception only replaces the point-tracking grasp check with a gripper-load and end-effector-lift test; flattened removes the reversibility-dependent threshold switch.
Appendix figures & tables17 assets
Supplementary material from the paper’s appendix.
Appendix
Task
Min
Median
Mean
Max
PickXTimes ( n=126 )
7
12
13.2
22
BinFill ( n=140 )
5
14
15.3
30
Appendix
Table 4: Lifting-clip length L in frames over the reported runs. The clip closes once a 0.03 m lift is observed, so L reflects the executed lift rather than a fixed horizon.
Layer
Simulation
Real robot
Linear(⋅,512)
2,369,024
2,368,000
Linear(512,128)
65,664
65,664
Linear(128,2)
258
258
Total
2,434,946
2,433,922
Appendix
Table 5: Parameter count of the verification head, the only trained component of AGM. Input widths are 4626 and 4624 respectively. The frozen SigLIP encoder (203.2M) and CoTracker3 tracker (25.4M) contribute no trainable parameters.
Symbol
Quantity
Value
Event localization
—
Closure entry width
0.030
—
Closure exit width
0.035
—
Empty-closure width
0.008
—
Hold-confirm
5 frames
Grasp verification
Appendix
Table 6: Complete simulation hyperparameter set, shared across all four tasks and all episodes.
Type
Total
Achieved
Failed
grasp
384
282
102
place-rev
278
221
0 57
place-irrev
126
0 87
0 39
Total
788
590
198
Appendix
Table 7: The 788 training events by achievement type. Failures constitute 25.1% of the corpus.
Controller
Type
Ach.
Fail
AGM (PickXTimes)
grasp
137
0 27
place-rev
165
00 0
Attempt-based (PickX)
grasp
0 40
0 72
place-rev
0 56
0 57
AGM (BinFill)
grasp
105
00 3
place-irrev
0 87
0 39
Appendix
Table 8: The same events by collection controller, over 61, 40, and 40 episodes respectively; both PickXTimes controllers draw on the same 100-seed training split, so their episode sets overlap.
Subset
n
Acc.
failed class
P
R
F1
All
138
91.3
87.5
70.0
77.8
grasp
0 69
92.8
100.0
68.8
81.5
place-rev
0 41
100.0
100.0
100.0
100.0
place-irrev
0 28
75.0
see text
Appendix
Table 9: Held-out performance of the measurement model at argmax , in percent. Grasp events are included because they supply auxiliary supervision, though deployment decides grasps by point tracking rather than by this head.
Subgoal source
PickXTimes
BinFill
AGM (verified writes)
100.0
84.0
Attempt-based (unverified writes)
32.0
82.0
Whole instruction (no pointer) †
36.0
40.0
Appendix
Table 10: Success rate (%) by what decides the current subgoal, on the official 50-episode test split. AGM and the attempt-based controller share one subgoal-conditioned checkpoint and differ only in whether the pointer waits for evidence. † : a policy trained for full-instruction conditioning rather than the released checkpoint of Table 1 , since a subgoal-conditioned checkpoint given an entire instruction is out of distribution and cannot serve as this control. AGM entries average two runs; the other two are single runs.
Figure 4: Simulation rollout of BinFill at N=5 under AGM. One panel per pointer transition, overlaid with the frame index, task instruction, action and state vectors, and the active subgoal. The eleven subgoals—five grasp–placement pairs and the terminal button press—are issued in order, and the pointer advances only on a verified achievement, so the episode terminates after exactly five cubes have entered the bin.
Figure 5: Simulation rollout of PickXTimes at N=5 under AGM, green cube. The same cube is grasped and placed five times, so the scene at every grasp is visually near-identical; the repetition index appears only in the subgoal supplied by the pointer.
Figure 6: Simulation rollout of PickXTimes at N=5 under AGM, red cube. The instruction, object, and target differ from Figure 5 , while thresholds, transition rules, and verifier weights are identical.
Hyperparameter
Value (identical across all four)
Training steps
40,000
Batch size
128
Optimizer
AdamW, gradient-norm clip 1.0
Learning rate
linear warmup to 5×10−5 over 10k steps, constant thereafter
EMA decay
0.999
Action horizon
20 frames
Appendix
Table 11: Policy fine-tuning configuration for the real-robot experiments. π0.5 uses quantile normalization, π0 uses z -score normalization, following each architecture’s default.
Quantity
Calibrated value
Measured width, fully open
0.0692 m
Measured width, holding 26 mm cube
0.0260 m
Measured width, empty closure
0.0031 m
Closure entry
0.0476 m
Closure exit
0.0541 m
Empty-closure threshold
0.0146 m
Appendix
Table 12: State machine calibration on the physical platform.
Check
Result
G matches annotated sequence
194 / 194
Pointer reaches terminal subgoal
194 / 194
Episode stop emitted
192 / 194
Median transition offset
7.5 frames (0.25 s)
Transitions within 15 frames (0.5 s)
96.5%
Appendix
Table 13: Offline replay of the deployment controller on all 194 demonstration episodes, consuming no robot time.
Figure 7: Real-robot rollout of BinFill at N=5 with a single-color instruction, a configuration absent from the demonstration corpus (Appendix D.2 ). Cubes of three colors share the workspace, so every grasp is color-selective while the pointer counts deliveries of the instructed color. The episode terminates after exactly five placements.
Figure 8: Real-robot rollout of PickXTimes at N=5 , green cube. The cube returns to the workspace between repetitions, so panels 1, 3, 5, 7 and 9 are the same scene; only the pointer distinguishes them.
Figure 9: Real-robot rollout of PickXTimes at N=5 , red cube, under the same verification head and thresholds as Figure 8 .
Component
Setting
Control / camera rate
30 Hz, 640×480 RGB ×2
State / action
7-D (6 joints + gripper width)
Action chunk
20 frames; replan every 10 frames or on subgoal switch
Policy input images
front + wrist, letterboxed 2242
Policy fine-tuning
40k steps, batch 128, lr 5×10−5 after 10k warmup, frozen encoder
Vision-language-action (VLA) models have driven rapid progress in robotic manipulation, demonstrating strong fine-grained control and promising performance on long-horizon tasks. However, many existing VLAs lack explicit access to interaction history, making them vulnerable to perceptual aliasing: similar current observations and robot states at different task stages may induce action ambiguity and lower success rate. Existing methods incorporate temporal or progress cues through feature conditioning, action-prior modification, or sampling guidance. However, methods that jointly fine-tune memory modules and the base VLA incur additional policy-training costs, motivating the separation of trainable history-conditioned steering from frozen base-policy refinement. We propose ActMem-VLA, a dual-expert handover architecture that augments a frozen, fine-tuned VLA with a memory plugin comprising a Mamba-based memory module and a lightweight PreAction Expert (PAE). Specifically, Mamba encodes executed-action history into memory that conditions PAE alongside current context. With these inputs, PAE steers task progression during early, high-noise denoising, then passes the partially denoised action to the frozen Action Expert (AE) to refine action details during the remaining low-noise steps. The fine-tuned base VLA remains frozen throughout training, while only the Mamba module and PAE are jointly optimized. On LIBERO-Mem, ActMem-VLA achieves 80.8% average success across all ten tasks, compared with 65.2% for π0.5 and 49.5% for MemoryVLA, while introducing only 3.45% additional parameters. Across four real-world tasks, it improves the average success rate over π0.5 by 28.8%.
Yaxin Zhao, Dianye Huang, Chenwei Wang +2
Harbin Institute of Technology, Harbin, China. · Medical Intelligence and Robotic Cognition (MIRoC) Lab, Department of Mechanical Engineering, The University of Hong Kong (HKU), Hong Kong SAR, China. · Department of Computing, The Hong Kong Polytechnic University, Hong Kong SAR, China
Memory is essential for long-horizon, partially observed robotic manipulation: a robot must remember which object was placed in a drawer, whose cup it moved, or how many action cycles have elapsed. Recent vision-language-action (VLA) models embed memory directly inside the policy, but benchmarks show no single in-policy mechanism covers all spatio-temporal dimensions, trailing oracle methods by a wide margin. We argue that memory type dictates where memory should reside: short-term perceptual memory (repetition, timing, retracing) belongs inside the policy, while long-term object memory (persistent spatial state, containment, event history) belongs outside as an explicit, readable record. We present Ledger, a harness that realizes this split over a single fine-tuned π0.5 policy by pairing an in-policy frame-sampling memory with an external spatio-temporal object memory, the ledger, built from a SAM3 tracker and a VLM captioner of the demonstration and read by an LLM planner that decides at step boundaries. On RoboMME, Ledger reaches the highest four-suite average among the evaluated methods, 64.3% (vs. 45.9% for the strongest prior method under identical evaluation), leading object reference (60.7% vs. 40.3%) and object permanence (86.7% vs. 56.2%) using a single set of weights. Choosing the memory source at runtime, from the instruction and the record, removes the need for a task-level router.
Visual-memory systems commonly retain or compress past observations. Robot control additionally requires interaction-derived state that no individual frame may explicitly represent, such as persistent identity relations, accumulated progress, or ordered procedures. We introduce Simple Agentic Robot Memory (SimpleARM), a training-free memory layer for frozen generalist robot policies. From the task instruction, SimpleARM specifies what to monitor; frozen perceptual tools maintain compact typed state online; structured access retrieves that state only when a proposed subgoal depends on history; and current-view grounding resolves recalled entities before execution. We evaluate SimpleARM on RoboMME, a benchmark of memory-dependent robot manipulation tasks that require history information no longer available in the current observation. Across all 16 tasks and three policy seeds, SimpleARM achieves 67.17% mean success, compared with 44.51% for the strongest non-oracle baseline. Matched ablations show mechanism specificity: removing relation, reference, progress, or route state produces large losses where the affected state is retrieved for control, while largely sparing other tasks. These results support a state-based view of robot memory: effective memory for control is not simply retained visual history, but compact task-relevant state derived from the interaction history.
Yuyou Zhang, Yunbei Zhang, Miao Li +4
Carnegie Mellon University · Tulane University · New York University +2