Vision-Language-Action (VLA) models have shown strong promise for general-purpose robotic manipulation, yet adapting them to new tasks and domains remains inefficient: existing methods often rely on parameter tuning, incurring substantial costs and risking catastrophic forgetting of previously learned tasks. To address this, we propose \textbf{Optimus-R}, a memory-centric VLA framework that formulates robotic adaptation as explicit query-skill memory tuning. Optimus-R introduces: (i) An \textbf{Inline Memory Interface for skill extraction}. It inserts learnable memory tokens into the VLA prefix stream, allowing the backbone to derive control-aware query and skill representations within the native action-conditioning pathway. (ii) A \textbf{Query-Skill Memory Bank for skill learning}. It externalizes skills into query prototypes for deciding \emph{what} to retrieve and skill values for specifying \emph{how} to act, supporting skill reuse and expansion with limited parameter updates. (iii) A lightweight \textbf{Bridge-and-Adapt mechanism for skill updating}. It aligns target-domain queries and skills with the existing memory space through a lightweight adapter and residual memory updates. Experiments on in-domain adaptation, cross-domain adaptation, and lifelong learning show that Optimus-R enables data-efficient skill learning while mitigating catastrophic forgetting.
Figures & tables
Figure 1: Comparison between parameter-centric adaptation and our memory-centric adaptation framework.
Figure 2: Overview of Optimus-R. Given an observation and language instruction, Optimus-R inserts learnable memory tokens into the VLA prefix stream, allowing the backbone to produce inline memory states aligned with the policy pathway. A dual-head memory readout derives a query embedding qt for retrieval and a skill embedding st for behavior representation, with st supervised by a skill decoder during training. The Query-Skill Memory Bank stores reusable latent skill slots as query prototypes and skill values {(pjq,pjs)} , which can be retrieved, updated, or expanded through clustering and residual memory updates. The retrieved skill is projected into VLA-compatible skill tokens and used to condition the flow policy for action-chunk generation.
Method
Data
LIBERO
CALVIN
RoboTwin 2.0
Spatial
Object
Goal
Long
Avg.
Avg. Len
Hard
DP ( Chi et al., 2025 )
100%
78.3
92.5
68.3
50.5
72.4
0.56
1.6
ACT ( Zhao et al., 2023 )
100%
–
–
–
–
–
–
3.5
DP3 ( Ze et al., 2024 )
100%
–
–
–
–
–
–
5.2
RDT ( Liu et al., 2024 )
100%
–
–
–
–
–
–
18.4
MemoryVLA ( Shi et al., 2025 )
100%
98.4
98.4
96.4
93.4
96.7
–
–
Table 1: Performance comparison on LIBERO ( Liu et al., 2023a ) , CALVIN ( Mees et al., 2022 ) , and RoboTwin 2.0 ( Chen et al., 2025 ) . We report the average success rate on each LIBERO task suite, the average completion length (Avg. Len) on CALVIN (ABC → D), and the success rate under the Hard setting on RoboTwin 2.0. Data denotes the training-data ratio. † denotes reproduced results.
Figure 3: Overview of Real-World tasks. We classify tasks into two categories (A and B) for lifelong setting: the model first learns categories A tasks and is then incrementally adapted to categories B tasks. Details are provided in the Appendix.
Method
Samples
R.O.
B.S.
P.P.
P.T.
P.C.
B.P.
M.C.
S.P.P.
B.Sq.
Avg
RDT ( Liu et al., 2024 )
80
20.0
40.0
35.0
50.0
35.0
30.0
20.0
5.0
0.0
26.1
OpenVLA ( Kim et al., 2024 )
80
15.0
35.0
45.0
55.0
40.0
20.0
15.0
20.0
10.0
28.3
OpenVLA-OFT ( Kim et al., 2025 )
80
65.0
75.0
80.0
70.0
55.0
45.0
50.0
20.0
30.0
54.4
π0 ( Black et al., 2024 )
80
75.0
75.0
80.0
85.0
45.0
55.0
50.0
30.0
20.0
57.2
π0.5 † ( Intelligence et al., 2025b )
20
25.0
20.0
35.0
25.0
30.0
25.0
10.0
5.0
0.0
19.4
Optimus-R
20
40.0
40.0
55.0
40.0
35.0
35.0
30.0
10.0
15.0
33.3
Table 2: Cross-Domain evaluation across 9 Real-World tasks. Samples denotes the number of training demonstrations per task. We report the success rate (%) for each task. Compared to the baselines, Optimus-R demonstrates superior performance across different data scales.
Figure 4: Real-world in-domain lifelong learning results. The model first learns Nold=4 tasks and is then incrementally adapted to Nnew=5 tasks. The first row reports the post-adaptation success rates on the previously learned tasks, together with the average forgetting rate, where lower forgetting indicates better retention. The second row reports the success rates on the newly learned tasks after incremental adaptation. Optimus-R achieves higher new-task performance while better preserving old-task performance.
Method
Total params
Added params
Trainable params
GPU hours
Inference Hz
Average SR (%) by data ratio
30%
50%
70%
100%
π0.5 (Full FT)
3.6B
–
3.6B
560
17
72.1
92.1
95.0
96.9
π0.5 (LoRA)
3.9B
0.3B
0.3B
128
12
71.5
86.9
91.2
96.5
OpenVLA-OFT (LoRA)
7.8B
0.3B
0.3B
320
10
66.4
83.5
89.1
96.4
Optimus-R
3.6B
7M
0.3B
86
14
88.6
95.6
96.6
97.8
Table 3: Adaptation cost and LIBERO performance. GPU hours correspond to the 100% data setting. Parameter counts are rounded; M and B denote millions and billions.
Model Variant
RoboTwin 2.0 (In-Domain)
Real-World 9 Tasks (Cross-Domain)
50% Data
100% Data
20 Demos
50 Demos
80 Demos
w/o Inline Memory
16.5
45.2
20.6
48.3
60.0
Coupled Query-Skill
18.2
48.0
22.8
50.0
62.8
w/o Prototype Residuals
17.0
46.5
17.8
46.7
58.3
w/o Stage-B Bridge
21.4
53.8
12.8
35.6
48.3
Optimus-R (Full)
21.6
55.0
33.3
61.7
72.2
Table 4: Ablation success rates (%) on RoboTwin 2.0 (Hard) and nine real-world tasks under different data budgets. Variant definitions are given in Sec. 4.6 .
Figure 5: Qualitative results of Optimus-R on real-world tasks. Each row shows a rollout sequence from left to right. Rows 1 and 3 correspond to learned tasks, where Optimus-R executes previously trained pick-and-place behaviors, including placing a red apple onto a plate and placing bottles onto a plate. Rows 2 and 4 show zero-shot task variations, where the model transfers the learned manipulation structure to new object-receptacle combinations, such as placing a green apple into a basket and placing blocks into a bowl.
Figure 6: Aligned queries colored by task. All 4,000 queries are shown; labels such as T34 denote task IDs. Prototype markers are omitted.
Appendix figures & tables11 assets
Supplementary material from the paper’s appendix.
Appendix
Parameter group / operation
Stage A
Stage B
Stage C
Memory tokens Emem
train
freeze
freeze
Poolers AttnPoolq,AttnPools
train
freeze
freeze
Encoders fq,fs
train
freeze
freeze
Skill Decoder Ds
train
freeze
freeze
Token Projector Tψ
train
freeze
freeze
Backbone Fθ
last layers
train
freeze
Appendix
Table 5: Stage-wise optimization and memory-bank operations. “Last layers” denotes the partial backbone training in Stage A. The optional Skill Adapter is a Stage-B component; it is fixed or inactive in Stage C.
Category
Hyperparameter
Value
Architecture
Memory tokens m
4
Query dimension dq
256
Skill dimension ds
256
Skill tokens ns
4
Memory bank / retrieval
Initial bank size K0
128
Retrieval count Kr
2
Appendix
Table 6: Reported default hyperparameters for Optimus-R.
Task
RDT
ACT
DP
DP3
π0
π0.5†
Optimus-R
Adjust Bottle
75%
23%
0%
3%
56%
75%
89%
Click Alarmclock
12%
4%
5%
14%
11%
44%
41%
Click Bell
9%
3%
0%
0%
3%
64%
53%
Grab Roller
43%
25%
0%
2%
80%
82%
94%
Move Playingcard Away
11%
0%
0%
3%
22%
32%
38%
Pick Diverse Bottles
0%
0%
0%
1%
6%
29%
36%
Appendix
Table 7: Performance comparison on RoboTwin 2.0. We report per-task success rates (SR) over 100 rollouts under Hard setting. † represents the result we reproduced.
Cat.
Task Family
Arm Mode
Objects
Targets
A
Reorientation
Single-arm
bottle, mug
-
Block Stacking
Single-arm
block
-
Place Object onto Plate
Single-arm
can, bowl, fruit
plate
Place Object onto Tablecloth
Single-arm
bowl, plate, cup
tablecloth
B
Put Objects into Container
Bimanual
bottle, can
basket
Bimanual Basket Placement
Bimanual
basket
tablecloth
Appendix
Table 8: Detailed configuration of the real-world tasks for lifelong learning evaluation. The tasks are divided into Base Tasks (Category A) and Novel Tasks (Category B), featuring a transition from single-arm to bimanual manipulation.
Figure 7: Qualitative rollouts on real-world Category-A tasks. We visualize representative trajectories for four previously learned manipulation tasks: Reorientation , Block Stacking , Place Object onto Plate , and Place Object onto Tablecloth . Across these tasks, Optimus-R consistently executes precise object-centric manipulation behaviors under diverse object and target configurations. These results illustrate that the learned query-skill memory can preserve reusable visuomotor primitives and support stable execution of previously acquired skills during subsequent adaptation.
Figure 8: Qualitative rollouts on real-world Category-B tasks under the lifelong adaptation setting. We show representative trajectories for five newly introduced tasks: Put Objects into Container , Bimanual Basket Placement , Multi-object Collection , Sequential Pick and Place , and Bimanual Coordination Sequential Task . These tasks require multi-object reasoning, long-horizon sequencing, and coordinated dual-arm control, posing a substantial distribution shift from the Category-A task set. Optimus-R successfully adapts to these novel behaviors by retrieving, updating, and expanding explicit query-skill memory entries, enabling new skill acquisition while mitigating interference with previously learned manipulation capabilities.
Variant
Average SR (%)
Optimus-R (Full)
97.8
w/o skill-aware clustering
95.8
w/o exemplar storage
96.3
Uniform downsampling
97.2
Appendix
Table 9: Memory-bank component ablations on LIBERO with 100% training data.
K0
Kr
m
Average SR (%)
64
2
4
97.2
128
2
4
97.8
256
2
4
97.8
128
1
4
97.0
128
4
8
97.4
128
4
1
96.8
Appendix
Table 10: LIBERO configuration sensitivity. Bold entries identify the default configuration. The last two rows jointly vary retrieval count and memory-token count.
Adaptation stage
K
New
Latency (ms)
Recall@5 (%)
After LIBERO-90
81
–
1.80
95.6
+ Spatial
84
3
1.80
95.4
+ Object
90
6
1.83
96.1
+ Goal
99
9
1.86
95.2
+ Long
109
10
1.90
94.7
Appendix
Table 11: Memory growth during sequential suite adaptation. Prototype increments are relative to the preceding stage; latency measures retrieval rather than full policy inference.
Evaluation stage
Goal SR (%)
Bowl-on-plate SR (%)
K
Zero-shot
8
68
16
After Goal adaptation
96
98
27
Appendix
Table 12: Spatial-to-Goal transfer. Bowl-on-plate denotes the Goal task put the bowl on the plate .
Checkpoint
T1
T2
T3
T4
T5
T6
T7
T8
T9
T10
After T1
100
–
–
–
–
–
–
–
–
–
After T2
98
100
–
–
–
–
–
–
–
–
After T3
98
100
100
–
–
–
–
–
–
–
After T4
100
98
100
98
–
–
–
–
–
–
After T5
100
100
94
100
98
–
–
–
–
–
After T6
96
100
96
98
94
98
–
–
–
–
Appendix
Table 13: Task-wise continual learning on LIBERO-Spatial. Entries are success rates (%); dashes denote tasks not yet introduced. T1 – T10 follow the sequence in Liu et al. (2026) .
Similar to the natural capabilities of humans to sequentially learn new tasks, robots with Vision-Language-Action (VLA) models should possess lifelong learning ability to learn a new task when deployed in open-world environments. However, most recently proposed lifelong learning models aim to effectively learn the current task (plasticity) or maintain high accuracy on previous tasks (stability), while the plasticity-stability trade-off remains largely unsolved in robotic manipulation models. To address this fundamental challenge, we propose a cache-efficient lifelong Vision-Language-Action learning framework for robotic manipulation (i.e., LifelongVLA), which alleviates the plasticity-stability trade-off with a dual-timescale adaptation mechanism while achieving low-cost robotic deployment with a cache-efficient replay strategy. More concretely, we propose a dual-timescale LoRA gating module to decompose VLA adaptation into two lightweight pathways: a short-term adapter for plasticity and a long-term adapter for stable consolidation. These pathways are integrated via a task-aware gate, enabling explicit control of the plasticity-stability trade-off. In the skill replay phase, a cache-efficient stochastic replay strategy is proposed to preserve more balanced retention signals without full-trajectory storage. Finally, experiments show that LifelongVLA outperforms existing baselines, demonstrating efficient skill expansion, robust retention of learned manipulation behaviors, and reduced reliance on retraining for real-world deployment on an xArm robot.
Yao He, Gan Sun, Wenqi Liang +2
South China University of Technology · University of Trento
Large-scale pretraining has made Vision-Language-Action (VLA) models promising foundations for generalist robot manipulation, yet adapting them to downstream tasks remains necessary. However, the common practice of full fine-tuning treats pretraining as initialization and can shift broad priors toward narrow training-distribution patterns. We propose PriorVLA, a novel framework that preserves pretrained priors and learns to leverage them for effective adaptation. PriorVLA keeps a frozen Prior Expert as a read-only prior source and trains an Adaptation Expert for downstream specialization. Expert Queries capture scene priors from the pretrained VLM and motor priors from the Prior Expert, integrating both into the Adaptation Expert to guide adaptation. Together, PriorVLA updates only 25% of the parameters updated by full fine-tuning. Across RoboTwin 2.0, LIBERO, and real-world tasks, PriorVLA achieves stronger overall performance than full fine-tuning and state-of-the-art VLA baselines, with the largest gains under out-of-distribution (OOD) and few-shot settings. PriorVLA improves over pi0.5 by 11 points on RoboTwin 2.0-Hard and achieves 99.1% average success on LIBERO. Across eight real-world tasks and two embodiments, PriorVLA reaches 81% in-distribution (ID) and 57% OOD success with standard data. With only 10 demonstrations per task, PriorVLA reaches 48% ID and 32% OOD success, surpassing pi0.5 by 24 and 22 points, respectively.
Xinyu Guo, Bin Xie, Wei Chai +4
Institute of Automation, Chinese Academy of Sciences · Dexmal · Nanjing University of Aeronautics and Astronautics +1
Vision-Language-Action (VLA) models have emerged as a promising paradigm for robotic manipulation by leveraging pre-trained vision-language representations. However, current VLA training methods suffer from two critical limitations: poor generalization to novel environments and low training efficiency requiring extensive demonstrations. We introduce Agentic-VLA, an agentic training framework that enables VLAs to efficiently adapt online through three key innovations: (1) Adaptive Reward Synthesis, which dynamically generates and adjusts reward functions based on the VLA's current capabilities and task complexity, decomposing complex tasks into learnable sub-goals for curriculum learning; (2) Language-Guided Exploration, where a critic model provides structured guidance for systematic exploration rather than random sampling; and (3) Experience Memory,which stores and retrieves task-relevant policy weights for warm-starting adaptation to similar tasks. We evaluate Agentic-VLA on the LIBERO benchmark, achieving substantial improvements: +12.3% on long-horizon tasks, +28.5% in 1-shot learning, and enabling cross-task transfer from 0% to 31.2% without task-specific demonstrations. Our framework also demonstrates 2.4x faster convergence compared to existing online adaptation methods. Beyond LIBERO, Agentic-VLA retains its advantage on the dual-arm RoboTwin 2.0 benchmark, including under its randomized Hard setting. These results establish Agentic-VLA as a significant step toward truly adaptive VLA systems capable of continuous learning in deployment.