Benchmarking and Enhancing Skill-Level Memory for Partially Observable Robotic Manipulation
Organizations: School of Information Science And Technology, University of Science and Technology of China · Shanghai Artificial Intelligence Laboratory · State Key Lab of Novel Software Technology, Nanjing University · Computer Science and Technology, Zhejiang University · Computer Science and Technology, Shanghai Jiaotong University
Abstract
Recent advances in robot learning have enabled manipulation policies to perform increasingly diverse tasks and generalize across environments. However, reliable execution often depends on hidden task states that cannot be determined from current observations alone, making interaction history essential. We introduce , a benchmark for evaluating manipulation memory under partial observability. HIDE comprises 15 tasks covering repetition counting, historical-state recall, and execution-progress tracking, with randomized initial configurations and decision points where similar observations require different actions depending on prior events. We further propose , a framework combining three complementary memory mechanisms to retain historical evidence and track execution state. Evaluations reveal substantial limitations in existing policies on HIDE, while memory augmentation improves task success in both simulation and real-world experiments. Individual mechanisms benefit some tasks but can degrade others; their combination achieves the highest average success rate on HIDE among the evaluated configurations. These findings highlight the importance of maintaining internal representations of hidden task states and matching memory design to task-specific information requirements.
Figures & tables
| Repetition Counting (RC) | Historical-State Recall (HSR) | Execution-Progress Tracking (EPT) | |||||||||||||||||
| Method | Avg. | Push Button | Stack Blocks | Change Channel | Stack Cups | Stack Blocks-N | Avg. | Weigh On/Off | Pour Back | Wipe Desk | Light Bulb | Reopen Drawer | Avg. | Search Drawer | Search Boxes | Search Cabinet | Lift & Check | Swap Pegs | Avg. |
| OpenVLA ( Kim et al., 2024 ) | 1.6 | 8 | 0 | 4 | 0 | 4 | 3.2 | 8 | 0 | 0 | 0 | 0 | 1.6 | 0 | 0 | 0 | 0 | 0 | 0.0 |
| OpenVLA-OFT Kim et al. (2025) | 13.9 | 12 | 0 | 28 | 0 | 0 | 8.0 | 36 | 0 | 68 | 0 | 32 | 27.2 | 0 | 4 | 0 | 28 | 0 | 6.4 |
| ( Black et al., 2024 ) | 22.1 | 40 | 0 | 12 | 0 | 0 | 10.4 | 40 | 24 | 92 | 0 | 24 | 36.0 | 12 | 28 | 36 | 24 | 0 | 20.0 |
| ( Physical Intelligence et al., 2025 ) | 14.4 | 16 | 0 | 4 | 4 | 4 | 5.6 | 20 | 8 | 76 | 0 | 20 | 24.8 | 4 | 8 | 20 | 32 | 0 | 12.8 |
| GR00T-N1.7 ( NVIDIA et al., 2025 ) | 26.4 | 12 | 0 | 12 | 4 | 4 | 6.4 | 96 | 0 | 84 | 0 | 68 | 49.6 | 28 | 16 | 40 | 24 | 8 | 23.2 |
| Method | Avg. | Push button N | Stack N cups (M total) | Clean desk | Search chip |
| 13 | 0 | 20 | 0 | 32 | |
| SAM2Act+ | 47 | 32 | 48 | 44 | 64 |
| SEEK | 89 | 76 | 100 | 80 | 100 |
Appendix figures & tables10 assets
Supplementary material from the paper’s appendix.
Appendix
| Task name | Language Template | Frames | Keyframes | # of Var. | Variation Type |
| push_button_times | “push the [color] button [N] times” | 144.3 | 7.8 | 50 | repetition count button color |
| stack_blocks | “stack [N] [color] blocks” | 372.8 | 16.1 | 60 | stack size color |
| change_channel_times | “turn the channel [direction] [N] times” | 244.5 | 12.2 | 6 | direction repetition count |
| stack_cups_new | “stack [N] cup(s) on top of the [color] cup” | 208.8 | 8.4 | 40 | base color cup count |
| stack_blocks_new | “put [N] [color] blocks into the rectangular container” | 283.5 | 11.8 | 60 | placement count color |
| weighing_on_off | “weigh the pepper and put it in another container” | 199.0 | 11.0 | 2 | starting container |
| Task | Language Template | Frames | Keyframes | # Var. | Variation Type |
| push_button_times | “push the [color] button [N] times” | 144.3 | 7.8 | 3 | repetition count randomized object placement |
| stack_cups | “stack [N] cups on the middle cup” | 372.8 | 16.1 | 4 | target stack count randomized object placement |
| wipe_desk_rubbish | “clean the desk” | 272.9 | 9.0 | 2 | rubbish location randomized object placement |
| lift_and_check | “lift the blocks and look for the white piece” | 277.3 | 14.3 | 3 | hidden-piece location randomized object placement |
| Hyperparameter | Stage 1 | Stage 2 |
| Number of GPUs | 8 | 8 |
| Batch size (frames) | 56 | 448 |
| Batch size (temporal sequences) | – | 64 |
| Frames per sequence | – | 7 |
| Gradient accumulation | 1 | 1 |
| Peak learning rate |
| Method | Avg. Success | Avg. Rank | Close Jar | Drag Stick | Insert Peg | Meat off Grill | Open Drawer | Place Cups | Place Wine | Push Buttons |
| Image-BC (CNN) ( Jang et al., 2022 ) | 1.3 | 14.08 | 0.0 | 0.0 | 0.0 | 0.0 | 4.0 | 0.0 | 0.0 | 0.0 |
| Image-BC (ViT) ( Jang et al., 2022 ) | 1.3 | 14.25 | 0.0 | 0.0 | 0.0 | 0.0 | 0.0 | 0.0 | 0.0 | 0.0 |
| C2F-ARM-BC ( James et al., 2022 ) | 20.1 | 13.03 | 24.0 | 24.0 | 4.0 | 20.0 | 20.0 | 0.0 | 8.0 | 72.0 |
| HiveFormer ( Guhur et al., 2022 ) | 45.3 | 11.06 | 52.0 | 76.0 | 0.0 | 100.0 | 52.0 | 0.0 | 80.0 | 84.0 |
| PolarNet ( Chen et al., 2023 ) | 46.4 | 10.19 | 36.0 | 92.0 | 4.0 | 100.0 | 84.0 | 0.0 | 40.0 | 96.0 |
| PerAct ( Shridhar et al., 2023 ) | 49.4 4.3 | 9.83 | 55.2 4.7 | 89.6 4.1 | 5.6 4.1 | 70.4 2.0 | 88.0 5.7 | 2.4 3.2 | 44.8 7.8 | 92.8 3.0 |
| Method | Clean | Average | MO-Color | RO-Color | MO-Texture | RO-Texture | MO-Size | RO-Size |
| PerAct | ||||||||
| RVT † | ||||||||
| RVT-2 | ||||||||
| SAM2Act | ||||||||
| w/o memory (Ours) | ||||||||
| w/ memory (Ours) |