TMem: Learning Test-Time Memory for Robotics
Organizations: Stanford University · NVIDIA · University of Michigan, Ann Arbor
Abstract
Memory-dependent robotic manipulation requires policies to use information that is no longer available in the current observation. Retaining history alone is insufficient: memory must preserve information that supports future actions. One challenge is whether a memory-free foundation model can learn to retain and use historical information from action demonstrations alone, without external memory support. We introduce TMem, a framework that develops this capability within a pretrained vision-language-action policy, without external reasoning models or memory-specific annotations. TMem uses test-time training to encode observation history into compact fast weights through online self-supervised updates, avoiding repeated processing of the full history. An observation-grounded interface extracts vision-language information for memory formation and supplies retrieved context to the action expert. Action supervision shapes what the memory learns to retain and use, while alternating memory-policy learning gives each component a fixed counterpart during optimization. Across 16 RoboMME tasks, TMem improves average success from 17.93% to 56.83% over the memory-free base policy and outperforms the recurrent-memory methods reported in the benchmark, while controlled profiling indicates at least 3x inference speedup over explicit methods. Project website: https://yzliu84.github.io/T2MEM-project/
Figures & tables
| Method | Counting | Permanence | Reference | Imitation | AVG | |||||||||||||
| Bin Fill | Pick Xtimes | Swing Xtimes | Stop Cube | Video Umsk | Button Umsk | Video UmskS | Button UmskS | Pick HighL | Video Repick | Video PlcBtn | Video PlcOrd | Move Cube | Insert Peg | Pattern Lock | Route Stick | |||
| Human Performance | 96.00 | 100.0 | 80.00 | 78.00 | 90.00 | 92.00 | 92.00 | 90.00 | 92.00 | 92.00 | 98.00 | 90.00 | 90.00 | 98.00 | 84.00 | 86.00 | 90.50 | |
| MME-VLA w/ Symbolic Memory | ||||||||||||||||||
| SimpleSG | Oracle | 85.78 | 99.78 | 100.0 | 44.67 | 33.11 | 22.00 | 15.56 | 15.56 | 44.00 | 27.78 | 31.33 | 26.00 | 87.33 | 10.00 | 95.33 | 55.11 | 49.58 |
| GroundSG | Oracle | 85.78 | 100.0 | 100.0 | 49.67 | 98.78 | 95.00 | 99.22 | 80.22 | 83.33 | 97.33 | 100.0 | 100.0 | 87.78 | 15.56 | 97.00 | 55.56 | 84.08 |
| SimpleSG | Gemini | 46.00 | 63.00 | 45.00 | 2.00 | 29.00 | 9.00 | 14.00 | 2.00 | 20.00 | 15.00 | 26.00 | 29.00 | 61.00 | 4.00 | 7.00 | 0.00 | 23.25 |
Appendix figures & tables8 assets
Supplementary material from the paper’s appendix.
Appendix
| Component | Reference configuration |
| VLM / action expert | 18 layers each; widths 2,048 / 1,024 |
| Interface | 16 tokens of width 1,024, initialized from ; rank-16 attention adapters |
| Fast-weight memory | 16 heads; per-head MLP , exact GeLU, biases |
| Learned initialization | : matrix standard deviation 0.02, zero biases |
| Read/write projections | Separate biased Q/K/V projections |
| Normalization / position | Q/K/V RMS normalization ( ); Q/K interleaved RoPE, base |
| Stage 1 | Stage 2A | Stage 2B | |
| Purpose | Memory-free adaptation | Memory learning | Memory-conditioned policy learning |
| Initialization | Pretrained | Stage-1 step 20,000; new memory | Preceding memory phase |
| Updated group | VLM LoRA, complete AE, action input/output and time projections | Interface, memory projections and initialization, step multipliers, gates | VLM LoRA, AE, action/time and proprioceptive projections |
| Budget | step 20,000 | 500 updates per cycle | 500 updates per cycle |
| Learning rate | ; gate | ||
| Effective batch | 64 frames (8 accumulation 8) | 32 sequences | 32 sequences |
| Intervention | Protocol |
| Disable writing | Zero effective updates from reset, including demonstration and execution. Retain , reads, and normal inference-call/RNG cadence. |
| Replace content | Build memory from the target demonstration, a conflicting donor demonstration, or no demonstration. Keep the target environment, instruction, and sampling seed fixed across conditions. Empty means learned , not all-zero weights. |
| Restrict write times | On a 66-frame prefix, permit all writes, only , or only . The write-timing protocol restricts writes to a 66-frame prefix, which truncates longer Hard demonstrations; the full-prefix condition therefore differs from the unrestricted setting in panel (a). |