Organizations: Xi’an Jiaotong University · The Hong Kong University of Science and Technology (Guangzhou) · Xbotics Embodied AI Community · The Chinese University of Hong Kong
Most vision-language-action (VLA) models rely on current observations and lose task-relevant evidence once it leaves view, limiting performance on long-horizon, memory-dependent tasks. Existing efforts incorporate compressed historical features or sparse visual keyframes. However, isolated snapshots can leave the policy uncertain about what changed during past interactions and which action should follow. To overcome this limitation, we propose EvoMem-VLA, which constructs state-evolution memory by explicitly encoding and retaining observed changes between historical states. These change representations preserve evidence of interaction outcomes, allowing the policy to track task progress beyond isolated snapshots. Specifically, we introduce conditional delta tokenization to encode ordered frame pairs into directional, source-conditioned delta tokens, each associated with its corresponding state evidence. A shared VLM backbone supports task-adaptive routing: normal long-horizon tasks follow a direct action route, whereas multi-stage tasks use a subtask route that generates an executable subtask as an additional input for action generation. With a single jointly trained policy for each simulation benchmark, EvoMem-VLA achieves success rates of 80.7% on RMBench, 82.0% on RoboMME and 83.8% across four real-world tasks spanning two robot embodiments. These results represent substantial improvements over the previous state of the art in all three evaluation settings.
Figures & tables
Figure 1 : Overview of EvoMem-VLA. EvoMem-VLA uses conditional delta tokenization to retain observed changes in state-evolution memory. (a) Motivation: sparse historical observations can leave important interaction changes unrecorded. (b) Comparison of memory-free VLAs, static-memory VLAs, and EvoMem-VLA. (c) State-of-the-art results on RMBench, RoboMME, and real-world tasks.
Figure 2 : Architecture of EvoMem-VLA. (a) A shared VLM performs action prediction, subtask generation, and keyframe selection from the current observation, state-evolution memory, instruction, and optional subtask history. Multi-stage tasks first generate an executable subtask as additional context for action generation. (b) The Delta Tokenizer maps an ordered frame pair to a directional, source-conditioned token. (c) State-Evolution Memory interleaves three recent delta-token groups with recent states and event-local delta tokens with selected keyframes.
Method
M(1)
M(n)
Overall Avg.
Observe & Pick Up
Rearrange Blocks
Put Back Block
Swap Blocks
Swap T
Avg.
Battery Try
Blocks Ranking Try
Cover Blocks
Press Button
Avg.
▼ Memory-free policies
DP ( Chi et al., 2025 )
1.0
0.0
0.0
11.0
20.0
6.4
10.0
10.0
0.0
0.0
5.0
5.8
ACT ( Zhao et al., 2023 )
1.0
29.0
0.0
2.0
2.0
6.8
19.0
0.0
0.0
0.0
4.8
5.9
π0.5 ( Physical Intelligence and others, 2025 )
9.0
13.0
11.0
24.0
15.0
14.4
16.0
6.0
0.0
0.0
5.5
10.4
X-VLA ( Zheng and others, 2025a )
9.0
13.0
18.0
16.0
3.0
11.8
26.0
1.0
2.0
0.0
7.3
9.8
Table 1 : RMBench results. Success rates (%) on the five M(1) and four M(n) tasks, together with group-wise and overall averages. EvoMem-VLA uses one policy jointly trained across all nine tasks. Bold : best; underlined : second best.
Method
Counting
Permanence
Reference
Imitation
Avg.
▼ Memory-free policies
π0.5 ( Physical Intelligence and others, 2025 )
28.8
17.0
17.2
8.8
17.9
▼ Memory-augmented policies
SimpleSG (QwenVL) ( Dai et al., 2026 )
44.6
19.6
25.2
26.6
29.0
GroundSG (QwenVL) ( Dai et al., 2026 )
38.0
39.3
31.6
21.9
32.7
TokenDrop (Modul) ( Dai et al., 2026 )
52.3
26.8
34.7
38.3
38.0
Table 2 : RoboMME results. Success rates (%) across four memory categories and the overall average. EvoMem-VLA uses one policy jointly trained across all 16 tasks. For baselines with multiple variants, we report the highest-performing non-oracle variant. Bold : best; underlined : second best.
Figure 3 : Real-world tasks and success rates. Four tasks across two robot embodiments. Results on the right report each task’s success rate (%) averaged across the two embodiments.
Figure 4 : Trajectory comparison on RMBench Press Button . Under the same episode initialization, EvoMem-VLA completes the task; the variant without delta tokens follows an incorrect trajectory, while the naive feature difference variant repeats actions. Red boxes highlight the next actions.
Method
M(1)
M(n)
Overall
EvoMem-VLA
84.0
76.5
80.7
w/ naive feature difference
52.0
41.0
47.1
w/o subtask route
84.0
33.0
61.3
w/o delta tokens
59.0
44.0
52.3
w/o delta tokens & subtask route
55.0
31.0
44.3
Table 3 : Component ablations on RMBench. Success rates (%) averaged over the five M(1) tasks, the four M(n) tasks, and all nine tasks. Bold : best; underlined : second best. All variants are implemented and evaluated within our framework under the same training and evaluation protocol.
Appendix figures & tables7 assets
Supplementary material from the paper’s appendix.
Appendix
Setting
Value
Feature extractor
DINOv3 ViT-B/16, frozen
Tokenizer input resolution
256×256
Encoder / decoder depth
12 / 12 layers
Delta-token dimension
768
Delta tokens per frame pair
1
Source–target interval
15 observation steps
Appendix
Table 4 : Delta Tokenizer implementation and adaptation settings for the simulation benchmarks.
Setting
RMBench
RoboMME
Camera streams
Head and two wrists
Front and wrist
Policy image resolution
224×224
224×224
Action dimension
14
8
Action-chunk horizon
50
20
Recent delta-token groups
3
3
Keyframe memory
Disabled
Enabled
Appendix
Table 5 : Policy configuration and optimization settings for the simulation benchmarks.
Figure 5 : Camera configurations of the two real-world robot platforms. Main-view and wrist cameras are marked.
Setting
Value
Demonstrations per task per embodiment
50
Demonstrations per embodiment
200
Policy initialization
StarVLA/VLAct_Qwen3_Pretrain
Policy fine-tuning
Full-parameter
Delta Tokenizer
Official pretrained checkpoint, frozen
Optimizer
AdamW
Appendix
Table 6 : Real-world training and deployment settings shared by both embodiments. Optimization settings follow EventVLA; initialization, action horizons, and memory construction follow our implementation.
Task
Language instruction
Varying conditions and success criteria
Lift Cup
Pick up the cup and put it back down the number of times shown on the card.
The requested count ranges from 3 to 10. The robot must complete exactly that many repetitions and return the cup; either too few or too many repetitions is a failure.
Cover Ducks
Cover the ducks with the cups from left to right, then uncover the pink, green, and blue ducks in that order.
Initial positions and color order vary. The robot must complete the covering stage and uncover all three ducks in the pink–green–blue order; any order violation is a failure.
Route Recall
Watch the order in which the blue boxes are pointed at, then place the cup into the boxes in the same order.
A human points to boxes in a varying sequence of 3–6 steps, including possible repeats. Success requires reproducing the entire demonstrated sequence, including repeated visits.
Put Back Cup
Move the cup to the center, return the arm to its home position, then pick up the cup and put it back in its original box.
The initial box varies. Success requires the complete sequence: move the cup to the center, return the arm home, and return the cup to its original box.
Appendix
Table 7 : Real-world task instructions and success criteria.
Benchmark
Numerical source
Coverage
RMBench
EventVLA, Appendix Table 8 ( Yang and others, 2026a )
All baseline task scores in Table 1 . DP, ACT, π0.5 , X-VLA, and Mem-0 are also reported in RMBench ( Chen and others, 2026a ) .
RoboMME
RoboMME benchmark report ( Dai et al., 2026 )
All baseline task scores underlying Table 2 , including the memory variants, past-action baseline, and MemER.
Appendix
Table 8 : Sources of simulation baseline scores. These are published results, not new baseline runs.
Variant
Change representation
Action routing
Subtask language loss
EvoMem-VLA
Conditional delta tokens
Task-adaptive
Enabled
w/ naive feature difference
Direct feature subtraction
Task-adaptive
Enabled
w/o subtask route
Conditional delta tokens
Direct only
Disabled
w/o delta tokens
None
Task-adaptive
Enabled
w/o delta tokens & subtask route
None
Direct only
Disabled
Appendix
Table 9 : Configurations of the RMBench ablations. Visual history and training settings are held fixed. Subtask supervision applies only to subtask-route samples.
Vision-language-action (VLA) models have driven rapid progress in robotic manipulation, demonstrating strong fine-grained control and promising performance on long-horizon tasks. However, many existing VLAs lack explicit access to interaction history, making them vulnerable to perceptual aliasing: similar current observations and robot states at different task stages may induce action ambiguity and lower success rate. Existing methods incorporate temporal or progress cues through feature conditioning, action-prior modification, or sampling guidance. However, methods that jointly fine-tune memory modules and the base VLA incur additional policy-training costs, motivating the separation of trainable history-conditioned steering from frozen base-policy refinement. We propose ActMem-VLA, a dual-expert handover architecture that augments a frozen, fine-tuned VLA with a memory plugin comprising a Mamba-based memory module and a lightweight PreAction Expert (PAE). Specifically, Mamba encodes executed-action history into memory that conditions PAE alongside current context. With these inputs, PAE steers task progression during early, high-noise denoising, then passes the partially denoised action to the frozen Action Expert (AE) to refine action details during the remaining low-noise steps. The fine-tuned base VLA remains frozen throughout training, while only the Mamba module and PAE are jointly optimized. On LIBERO-Mem, ActMem-VLA achieves 80.8% average success across all ten tasks, compared with 65.2% for π0.5 and 49.5% for MemoryVLA, while introducing only 3.45% additional parameters. Across four real-world tasks, it improves the average success rate over π0.5 by 28.8%.
Yaxin Zhao, Dianye Huang, Chenwei Wang +2
Harbin Institute of Technology, Harbin, China. · Medical Intelligence and Robotic Cognition (MIRoC) Lab, Department of Mechanical Engineering, The University of Hong Kong (HKU), Hong Kong SAR, China. · Department of Computing, The Hong Kong Polytechnic University, Hong Kong SAR, China
Memory remains a critical bottleneck for long-horizon robotic manipulation, as standard Vision-Language-Action (VLA) policies often fail when task-relevant cues become occluded or unobservable over time. While existing memory-augmented methods utilize historical context, they either suffer from severe information bottlenecks, incur high latency via decoupled dual systems, or rely on unselective buffers that accumulate massive visual redundancies. To address these limitations, we introduce EventVLA, an end-to-end framework founded on the concept of sparse visual evidence memory that comprises two core components: foundational visual anchors to retain initial and short-term contexts, and a dynamic Keyframe Evidence Memory (KEM) module. Specifically, KEM directly predicts future keyframe probabilities from the VLA's latent embeddings to autonomously capture and store sparse, task-critical visual events. This foresight-driven mechanism empowers the policy to dynamically evaluate the future causal utility of current observations, preserving transient visual evidence before it becomes unobservable. Furthermore, we propose RoboTwin-MeM, a diagnostic benchmark specifically designed to evaluate non-Markovian manipulation tasks with interactive visual evidence. Extensive evaluations show that across 17 memory-requiring simulation tasks and 4 real-world bimanual tasks, EventVLA achieves an average success rate improvement of +40% over state-of-the-art memory-augmented VLAs.
Ganlin Yang, Zhangzheng Tu, Yuqiang Yang +10
University of Science and Technology of China · Shanghai AI Laboratory · Dalian University of Technology +5
How can pretrained Vision-Language-Action (VLA) models retain long-horizon visual histories with high-frequency updates without sacrificing efficiency? Existing approaches rely on external memory management, which restrains either the memory horizon or the reactiveness of pretrained policies. To this end, we present NativeMEM, a VLA policy that features long-term and real-time updated memory. At its core is an efficient memory encoding scheme, Native Memory Compression, which repurposes the VLA's own vision encoder to compress each historical frame from each camera view into a single token. Appended to the input sequence, these memory tokens enable the pretrained VLA to attend over long-term history with negligible latency overhead, requiring neither an external planner nor a freshly initialized memory module. To align the memory tokens with the pretrained policy, we first develop a generic memory tokenizer under the supervision of a frozen VLA on memory-demanding data, and then unfreeze the VLA for task-specific fine-tuning. NativeMEM consistently outperforms prior methods, boosting success rates from 32.4% to 84.0% in simulation and up to 98.7% on real robots, while maintaining low inference latency and GPU memory usage. Notably, NativeMEM exhibits high data efficiency by achieving competitive results with prior arts using only 20% of the training data.
Ziye Wang, Modi Shi, Chaojun Ni +5
The University of Hong Kong · Beihang University · Peking University +2