OCC4M: Object-Centric 4D Memory for Spatiotemporal Reasoning in Long-Horizon Manipulation
Organizations: Harvard University · ETH Zürich
Abstract
Long-horizon manipulation often requires reasoning about state absent from the current view, such as a vanished object's location, temporal identity, or the contents of a shuffled container. We present OCC4M ("Occam"), an object-centric 4D memory that maintains persistent tracks in a shared world frame and explicitly represents temporal, motion, and containment relations. A vision-language model (VLM) queries this structured memory to select actionable targets for history-free low-level execution. Across seven simulation conditions and 350 episodes, OCC4M achieves 96.6% memory success and 88.9% end-to-end success, versus 54.6% and 57.7% for FrameSamp, a raw-history VLM baseline using Gemini 3.7 Flash with the complete observation history and the same executor. In a controlled viewpoint-transfer test, OCC4M maintains 100% memory and 98% end-to-end success after a viewpoint change, while full-history FrameSamp falls to near-zero success. On 20 fixed-camera Franka episodes, OCC4M reaches 85% joint memory accuracy, versus at most 30% for FrameSamp across context sizes from to the complete history, and completes 45% of full two-stage tasks. These results support explicit object-centric memory for persistent spatiotemporal reasoning in long-horizon manipulation. Qualitative videos are available at https://occ4m-sup.github.io/occ4m-supplementary/.
Figures & tables
| Capability | Required decision | |
|---|---|---|
| T1 | Spatial persistence | Place the red cube at the vanished green cube’s last location. |
| T1c | Viewpoint transfer | After the target vanishes and the base moves , express its location in the current view. |
| T2 | Temporal identity | Place the blue cube at the location of the marker that appeared first, and the red cube at the location of the marker that appeared second. |
| T3 | Event history | Three cubes move sequentially; pick the second mover and place it at the first mover’s pre-motion pose. |
| T4 | Relational persistence | After identical covers are shuffled, pick the cover containing the red cube. |
| T5 | Cross-view association | Retrieve a vanished sign location after driving to a second table and back. |
| Method | T1 | T1c | T2 | T3 | T4 | T5 | T6 | Mean |
| A. Memory success | ||||||||
| OCC4M | 100 | 100 | 92 | 98 | 88 | 98 | 100 | 96.6 |
| OCC4M + oracle actions | 100 | 100 | 92 | 98 | 88 | 98 | 100 | 96.6 |
| OCC4M + oracle memory | 100 | 100 | 100 | 100 | 100 | 100 | 100 | 100.0 |
| OCC4M + oracle memory+plan | 100 | 100 | 100 | 100 | 100 | 100 | 100 | 100.0 |
| FrameSamp ( , Gemini 3.7) | 100 | 6 | 72 | 32 | 54 | 100 | 2 | 52.3 |
| Condition | FrameSamp | OCC4M |
|---|---|---|
| T1: historical-frame grounding ( ) | 100 / 96 | 100 / 94 |
| T1b: current frame, fixed viewpoint ( ) | 100 / 92 | 100 / 94 |
| T1c: current frame + base motion ( ) | 6 / 6 | 100 / 98 |
| Memory accuracy | Physical execution | |||||
|---|---|---|---|---|---|---|
| Method | S1 | S2 | Both | S1 | S2 | E2E |
| OCC4M | 85% | 95% | 85% | 75% | 55% | 45% |
| FrameSamp ( ) | 30% | 95% | 30% | – | – | – |
| FrameSamp ( ) | 25% | 100% | 25% | – | – | – |
| FrameSamp (all frames) | 25% | 90% | 20% | – | – | – |
Appendix figures & tables9 assets
Supplementary material from the paper’s appendix.
Appendix
| Varying demonstrations per task (20k steps) | ||||
|---|---|---|---|---|
| Demonstrations | 80 | 250 | 400 | – |
| Mean success | 8.3 | 5.0 | 9.3 | – |
| Varying training duration (250 demonstrations/task) | ||||
| Training steps | 10k | 20k | 30k | 40k |
| Mean success | 1.7 | 5.0 | 3.3 | 3.3 |
| Task | Pick target (GT) | Place target (GT) | Additional constraints |
|---|---|---|---|
| T1 | red cube (current) | remembered green-cube location | N/A |
| T1c | red cube (current) | same remembered location, expressed after base motion | place prediction uses final frame |
| T2 | blue cube, red cube | blue first marker; red second marker | labels and order correct |
| T3 | second cube that moved (current) | pre-motion location of first mover | pick is second mover |
| T4 | current correct grey cover | N/A (pick-only) | correct cover among three, pick prediction uses final frame |
| T5 | query-color cube on Table B | remembered query-color sign location on Table A | queried color correct |