EvoMem-VLA: State-Evolution Memory for Long-Horizon Robot Manipulation
Organizations: Xi’an Jiaotong University · The Hong Kong University of Science and Technology (Guangzhou) · Xbotics Embodied AI Community · The Chinese University of Hong Kong
Abstract
Most vision-language-action (VLA) models rely on current observations and lose task-relevant evidence once it leaves view, limiting performance on long-horizon, memory-dependent tasks. Existing efforts incorporate compressed historical features or sparse visual keyframes. However, isolated snapshots can leave the policy uncertain about what changed during past interactions and which action should follow. To overcome this limitation, we propose EvoMem-VLA, which constructs state-evolution memory by explicitly encoding and retaining observed changes between historical states. These change representations preserve evidence of interaction outcomes, allowing the policy to track task progress beyond isolated snapshots. Specifically, we introduce conditional delta tokenization to encode ordered frame pairs into directional, source-conditioned delta tokens, each associated with its corresponding state evidence. A shared VLM backbone supports task-adaptive routing: normal long-horizon tasks follow a direct action route, whereas multi-stage tasks use a subtask route that generates an executable subtask as an additional input for action generation. With a single jointly trained policy for each simulation benchmark, EvoMem-VLA achieves success rates of 80.7% on RMBench, 82.0% on RoboMME and 83.8% across four real-world tasks spanning two robot embodiments. These results represent substantial improvements over the previous state of the art in all three evaluation settings.
Figures & tables
| Method | Overall Avg. | |||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Observe & Pick Up | Rearrange Blocks | Put Back Block | Swap Blocks | Swap T | Avg. | Battery Try | Blocks Ranking Try | Cover Blocks | Press Button | Avg. | ||
| Memory-free policies | ||||||||||||
| DP ( Chi et al., 2025 ) | 1.0 | 0.0 | 0.0 | 11.0 | 20.0 | 6.4 | 10.0 | 10.0 | 0.0 | 0.0 | 5.0 | 5.8 |
| ACT ( Zhao et al., 2023 ) | 1.0 | 29.0 | 0.0 | 2.0 | 2.0 | 6.8 | 19.0 | 0.0 | 0.0 | 0.0 | 4.8 | 5.9 |
| ( Physical Intelligence and others, 2025 ) | 9.0 | 13.0 | 11.0 | 24.0 | 15.0 | 14.4 | 16.0 | 6.0 | 0.0 | 0.0 | 5.5 | 10.4 |
| X-VLA ( Zheng and others, 2025a ) | 9.0 | 13.0 | 18.0 | 16.0 | 3.0 | 11.8 | 26.0 | 1.0 | 2.0 | 0.0 | 7.3 | 9.8 |
| Method | Counting | Permanence | Reference | Imitation | Avg. |
|---|---|---|---|---|---|
| Memory-free policies | |||||
| ( Physical Intelligence and others, 2025 ) | 28.8 | 17.0 | 17.2 | 8.8 | 17.9 |
| Memory-augmented policies | |||||
| SimpleSG (QwenVL) ( Dai et al., 2026 ) | 44.6 | 19.6 | 25.2 | 26.6 | 29.0 |
| GroundSG (QwenVL) ( Dai et al., 2026 ) | 38.0 | 39.3 | 31.6 | 21.9 | 32.7 |
| TokenDrop (Modul) ( Dai et al., 2026 ) | 52.3 | 26.8 | 34.7 | 38.3 | 38.0 |
| Method | Overall | ||
|---|---|---|---|
| EvoMem-VLA | 84.0 | 76.5 | 80.7 |
| w/ naive feature difference | 52.0 | 41.0 | 47.1 |
| w/o subtask route | 84.0 | 33.0 | 61.3 |
| w/o delta tokens | 59.0 | 44.0 | 52.3 |
| w/o delta tokens & subtask route | 55.0 | 31.0 | 44.3 |
Appendix figures & tables7 assets
Supplementary material from the paper’s appendix.
Appendix
| Setting | Value |
|---|---|
| Feature extractor | DINOv3 ViT-B/16, frozen |
| Tokenizer input resolution | |
| Encoder / decoder depth | 12 / 12 layers |
| Delta-token dimension | 768 |
| Delta tokens per frame pair | 1 |
| Source–target interval | 15 observation steps |
| Setting | RMBench | RoboMME |
|---|---|---|
| Camera streams | Head and two wrists | Front and wrist |
| Policy image resolution | ||
| Action dimension | 14 | 8 |
| Action-chunk horizon | 50 | 20 |
| Recent delta-token groups | 3 | 3 |
| Keyframe memory | Disabled | Enabled |
| Setting | Value |
|---|---|
| Demonstrations per task per embodiment | 50 |
| Demonstrations per embodiment | 200 |
| Policy initialization | StarVLA/VLAct_Qwen3_Pretrain |
| Policy fine-tuning | Full-parameter |
| Delta Tokenizer | Official pretrained checkpoint, frozen |
| Optimizer | AdamW |
| Task | Language instruction | Varying conditions and success criteria |
|---|---|---|
| Lift Cup | Pick up the cup and put it back down the number of times shown on the card. | The requested count ranges from 3 to 10. The robot must complete exactly that many repetitions and return the cup; either too few or too many repetitions is a failure. |
| Cover Ducks | Cover the ducks with the cups from left to right, then uncover the pink, green, and blue ducks in that order. | Initial positions and color order vary. The robot must complete the covering stage and uncover all three ducks in the pink–green–blue order; any order violation is a failure. |
| Route Recall | Watch the order in which the blue boxes are pointed at, then place the cup into the boxes in the same order. | A human points to boxes in a varying sequence of 3–6 steps, including possible repeats. Success requires reproducing the entire demonstrated sequence, including repeated visits. |
| Put Back Cup | Move the cup to the center, return the arm to its home position, then pick up the cup and put it back in its original box. | The initial box varies. Success requires the complete sequence: move the cup to the center, return the arm home, and return the cup to its original box. |
| Benchmark | Numerical source | Coverage |
|---|---|---|
| RMBench | EventVLA, Appendix Table 8 ( Yang and others, 2026a ) | All baseline task scores in Table 1 . DP, ACT, , X-VLA, and Mem-0 are also reported in RMBench ( Chen and others, 2026a ) . |
| RoboMME | RoboMME benchmark report ( Dai et al., 2026 ) | All baseline task scores underlying Table 2 , including the memory variants, past-action baseline, and MemER. |
| Variant | Change representation | Action routing | Subtask language loss |
|---|---|---|---|
| EvoMem-VLA | Conditional delta tokens | Task-adaptive | Enabled |
| w/ naive feature difference | Direct feature subtraction | Task-adaptive | Enabled |
| w/o subtask route | Conditional delta tokens | Direct only | Disabled |
| w/o delta tokens | None | Task-adaptive | Enabled |
| w/o delta tokens & subtask route | None | Direct only | Disabled |