AGM: Achievement-Grounded Memory for Closed-Loop Agents with Frozen VLA Policies
Organizations: Harbin Institute of Technology · The University of Sydney · StellarEdge AI
Abstract
Frozen vision-language-action (VLA) policies offer broad manipulation skills but execute open-loop action chunks without tracking task progress, so the agent cannot reliably decide whether to continue, retry, or terminate. External memory is a natural remedy, yet we find it can be harmful when attempted actions are recorded as completed progress: transient execution failures become persistent task-state errors, and such a memory can underperform no progress memory at all. We propose Achievement-Grounded Memory (AGM), a lightweight closed-loop framework for frozen VLA policies. AGM represents a task as a static subgoal sequence with a dynamic progress pointer and advances the pointer on physically verified achievement rather than on attempts. Proprioceptive gripper-load cues decide when to verify; coherent point tracking verifies grasps, and language-conditioned cross-view comparison, read by a single trained 2.43M-parameter verification head, verifies placements. The policy, tracker, and encoder remain frozen, the head is the only trained component, and deployment needs no auxiliary vision-language model. On the RoboMME Counting benchmark, AGM reaches 100.0% on PickXTimes and 84.0% on BinFill, surpassing the strongest memory-augmented baseline by 7.7 points on the four-task average, and the gains carry over to a physical robot, where AGM reaches 100.0% and 82.0%. These results suggest that reliable embodied memory depends more on disciplined state updates than on memory capacity.
Figures & tables
| Success Rate (%) | Params. | ||||||
|---|---|---|---|---|---|---|---|
| Method | PickXTimes | BinFill | SwingXTimes | StopCube | Avg. | External | Trainable |
| Frozen | 42.89 | 30.00 | 35.56 | 6.67 | 28.78 | 0 | 0 |
| w/ past actions | 58.33 | 26.67 | 26.67 | 4.67 | 29.09 | 0 | 3.35B |
| SAM2Act+ ‡ | 76.00 | 40.00 | 25.33 | 0.00 | 35.33 | – | n/r |
| SimpleSG + Qwen3-VL-4B | 95.30 | 77.60 | 5.10 | 0.40 | 44.60 | 4B | LoRA ( n/r ) |
| SimpleSG + Gemini-2.5-Pro | 63.00 | 46.00 | 45.00 | 2.00 | 39.00 | API | 0 |
| PickXTimes | BinFill | ||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|
| Method | |||||||||||
| Sim | 12.5 | 0.0 | 0.0 | 0.0 | 0.0 | 0.0 | 0.0 | 0.0 | 0.0 | 0.0 | |
| +AGM | 50.0 | 40.0 | 33.3 | 50.0 | 50.0 | 30.0 | 10.0 | 0.0 | 0.0 | 0.0 | |
| 100.0 | 30.0 | 16.7 | 0.0 | 0.0 | 90.0 | 20.0 | 18.8 | 12.5 | 0.0 | ||
| +AGM | 100.0 | 100.0 | 100.0 | 100.0 | 100.0 | 100.0 | 95.0 | 84.4 | 62.5 | 66.7 | |
| Real | 20.0 | 0.0 | 0.0 | 0.0 | 0.0 | 0.0 | 0.0 | 0.0 | 0.0 | 0.0 | |
| Grasp verification | Thresholds | PickXTimes | BinFill |
|---|---|---|---|
| CoTracker (full) | reversibility-aware | 100.0 | 84.0 |
| Proprioception only | reversibility-aware | 98.0 | 78.0 |
| CoTracker | flattened | 98.0 | 74.0 |
Appendix figures & tables17 assets
Supplementary material from the paper’s appendix.
Appendix
| Task | Min | Median | Mean | Max |
|---|---|---|---|---|
| PickXTimes ( ) | 7 | 12 | 13.2 | 22 |
| BinFill ( ) | 5 | 14 | 15.3 | 30 |
| Layer | Simulation | Real robot |
|---|---|---|
| Total |
| Symbol | Quantity | Value |
| Event localization | ||
| — | Closure entry width | |
| — | Closure exit width | |
| — | Empty-closure width | |
| — | Hold-confirm | 5 frames |
| Grasp verification | ||
| Type | Total | Achieved | Failed |
|---|---|---|---|
| grasp | 384 | 282 | 102 |
| place-rev | 278 | 221 | 0 57 |
| place-irrev | 126 | 0 87 | 0 39 |
| Total | 788 | 590 | 198 |
| Controller | Type | Ach. | Fail |
|---|---|---|---|
| AGM (PickXTimes) | grasp | 137 | 0 27 |
| place-rev | 165 | 00 0 | |
| Attempt-based (PickX) | grasp | 0 40 | 0 72 |
| place-rev | 0 56 | 0 57 | |
| AGM (BinFill) | grasp | 105 | 00 3 |
| place-irrev | 0 87 | 0 39 |
| Subset | Acc. | failed class | |||
|---|---|---|---|---|---|
| P | R | F1 | |||
| All | 138 | 91.3 | 87.5 | 70.0 | 77.8 |
| grasp | 0 69 | 92.8 | 100.0 | 68.8 | 81.5 |
| place-rev | 0 41 | 100.0 | 100.0 | 100.0 | 100.0 |
| place-irrev | 0 28 | 75.0 | see text | ||
| Subgoal source | PickXTimes | BinFill |
|---|---|---|
| AGM (verified writes) | 100.0 | 84.0 |
| Attempt-based (unverified writes) | 32.0 | 82.0 |
| Whole instruction (no pointer) † | 36.0 | 40.0 |
| Hyperparameter | Value (identical across all four) |
|---|---|
| Training steps | 40,000 |
| Batch size | 128 |
| Optimizer | AdamW, gradient-norm clip |
| Learning rate | linear warmup to over 10k steps, constant thereafter |
| EMA decay | |
| Action horizon | 20 frames |
| Quantity | Calibrated value |
|---|---|
| Measured width, fully open | m |
| Measured width, holding 26 mm cube | m |
| Measured width, empty closure | m |
| Closure entry | m |
| Closure exit | m |
| Empty-closure threshold | m |
| Check | Result |
|---|---|
| matches annotated sequence | 194 / 194 |
| Pointer reaches terminal subgoal | 194 / 194 |
| Episode stop emitted | 192 / 194 |
| Median transition offset | 7.5 frames (0.25 s) |
| Transitions within 15 frames (0.5 s) | 96.5% |
| Component | Setting |
|---|---|
| Control / camera rate | 30 Hz, RGB |
| State / action | 7-D (6 joints gripper width) |
| Action chunk | 20 frames; replan every 10 frames or on subgoal switch |
| Policy input images | front wrist, letterboxed |
| Policy fine-tuning | 40k steps, batch 128, lr after 10k warmup, frozen encoder |
| Grasp verification | lift m (proprioceptive); px, grid, px half-width |