Vision-Language-Action (VLA) models have shown strong promise for general-purpose robotic manipulation, yet adapting them to new tasks and domains remains inefficient: existing methods often rely on parameter tuning, incurring substantial costs and risking catastrophic forgetting of previously learned tasks. To address this, we propose \textbf{Optimus-R}, a memory-centric VLA framework that formulates robotic adaptation as explicit query-skill memory tuning. Optimus-R introduces: (i) An \textbf{Inline Memory Interface for skill extraction}. It inserts learnable memory tokens into the VLA prefix stream, allowing the backbone to derive control-aware query and skill representations within the native action-conditioning pathway. (ii) A \textbf{Query-Skill Memory Bank for skill learning}. It externalizes skills into query prototypes for deciding \emph{what} to retrieve and skill values for specifying \emph{how} to act, supporting skill reuse and expansion with limited parameter updates. (iii) A lightweight \textbf{Bridge-and-Adapt mechanism for skill updating}. It aligns target-domain queries and skills with the existing memory space through a lightweight adapter and residual memory updates. Experiments on in-domain adaptation, cross-domain adaptation, and lifelong learning show that Optimus-R enables data-efficient skill learning while mitigating catastrophic forgetting.
Figures & tables
Figure 1: Comparison between parameter-centric adaptation and our memory-centric adaptation framework.
Figure 2: Overview of Optimus-R. Given an observation and language instruction, Optimus-R inserts learnable memory tokens into the VLA prefix stream, allowing the backbone to produce inline memory states aligned with the policy pathway. A dual-head memory readout derives a query embedding qt for retrieval and a skill embedding st for behavior representation, with st supervised by a skill decoder during training. The Query-Skill Memory Bank stores reusable latent skill slots as query prototypes and skill values {(pjq,pjs)} , which can be retrieved, updated, or expanded through clustering and residual memory updates. The retrieved skill is projected into VLA-compatible skill tokens and used to condition the flow policy for action-chunk generation.
Method
Data
LIBERO
CALVIN
RoboTwin 2.0
Spatial
Object
Goal
Long
Avg.
Avg. Len
Hard
DP ( Chi et al., 2025 )
100%
78.3
92.5
68.3
50.5
72.4
0.56
1.6
ACT ( Zhao et al., 2023 )
100%
–
–
–
–
–
–
3.5
DP3 ( Ze et al., 2024 )
100%
–
–
–
–
–
–
5.2
RDT ( Liu et al., 2024 )
100%
–
–
–
–
–
–
18.4
MemoryVLA ( Shi et al., 2025 )
100%
98.4
98.4
96.4
93.4
96.7
–
–
Table 1: Performance comparison on LIBERO ( Liu et al., 2023a ) , CALVIN ( Mees et al., 2022 ) , and RoboTwin 2.0 ( Chen et al., 2025 ) . We report the average success rate on each LIBERO task suite, the average completion length (Avg. Len) on CALVIN (ABC → D), and the success rate under the Hard setting on RoboTwin 2.0. Data denotes the training-data ratio. † denotes reproduced results.
Figure 3: Overview of Real-World tasks. We classify tasks into two categories (A and B) for lifelong setting: the model first learns categories A tasks and is then incrementally adapted to categories B tasks. Details are provided in the Appendix.
Method
Samples
R.O.
B.S.
P.P.
P.T.
P.C.
B.P.
M.C.
S.P.P.
B.Sq.
Avg
RDT ( Liu et al., 2024 )
80
20.0
40.0
35.0
50.0
35.0
30.0
20.0
5.0
0.0
26.1
OpenVLA ( Kim et al., 2024 )
80
15.0
35.0
45.0
55.0
40.0
20.0
15.0
20.0
10.0
28.3
OpenVLA-OFT ( Kim et al., 2025 )
80
65.0
75.0
80.0
70.0
55.0
45.0
50.0
20.0
30.0
54.4
π0 ( Black et al., 2024 )
80
75.0
75.0
80.0
85.0
45.0
55.0
50.0
30.0
20.0
57.2
π0.5 † ( Intelligence et al., 2025b )
20
25.0
20.0
35.0
25.0
30.0
25.0
10.0
5.0
0.0
19.4
Optimus-R
20
40.0
40.0
55.0
40.0
35.0
35.0
30.0
10.0
15.0
33.3
Table 2: Cross-Domain evaluation across 9 Real-World tasks. Samples denotes the number of training demonstrations per task. We report the success rate (%) for each task. Compared to the baselines, Optimus-R demonstrates superior performance across different data scales.
Figure 4: Real-world in-domain lifelong learning results. The model first learns Nold=4 tasks and is then incrementally adapted to Nnew=5 tasks. The first row reports the post-adaptation success rates on the previously learned tasks, together with the average forgetting rate, where lower forgetting indicates better retention. The second row reports the success rates on the newly learned tasks after incremental adaptation. Optimus-R achieves higher new-task performance while better preserving old-task performance.
Method
Total params
Added params
Trainable params
GPU hours
Inference Hz
Average SR (%) by data ratio
30%
50%
70%
100%
π0.5 (Full FT)
3.6B
–
3.6B
560
17
72.1
92.1
95.0
96.9
π0.5 (LoRA)
3.9B
0.3B
0.3B
128
12
71.5
86.9
91.2
96.5
OpenVLA-OFT (LoRA)
7.8B
0.3B
0.3B
320
10
66.4
83.5
89.1
96.4
Optimus-R
3.6B
7M
0.3B
86
14
88.6
95.6
96.6
97.8
Table 3: Adaptation cost and LIBERO performance. GPU hours correspond to the 100% data setting. Parameter counts are rounded; M and B denote millions and billions.
Model Variant
RoboTwin 2.0 (In-Domain)
Real-World 9 Tasks (Cross-Domain)
50% Data
100% Data
20 Demos
50 Demos
80 Demos
w/o Inline Memory
16.5
45.2
20.6
48.3
60.0
Coupled Query-Skill
18.2
48.0
22.8
50.0
62.8
w/o Prototype Residuals
17.0
46.5
17.8
46.7
58.3
w/o Stage-B Bridge
21.4
53.8
12.8
35.6
48.3
Optimus-R (Full)
21.6
55.0
33.3
61.7
72.2
Table 4: Ablation success rates (%) on RoboTwin 2.0 (Hard) and nine real-world tasks under different data budgets. Variant definitions are given in Sec. 4.6 .
Figure 5: Qualitative results of Optimus-R on real-world tasks. Each row shows a rollout sequence from left to right. Rows 1 and 3 correspond to learned tasks, where Optimus-R executes previously trained pick-and-place behaviors, including placing a red apple onto a plate and placing bottles onto a plate. Rows 2 and 4 show zero-shot task variations, where the model transfers the learned manipulation structure to new object-receptacle combinations, such as placing a green apple into a basket and placing blocks into a bowl.
Figure 6: Aligned queries colored by task. All 4,000 queries are shown; labels such as T34 denote task IDs. Prototype markers are omitted.
Appendix figures & tables11 assets
Supplementary material from the paper’s appendix.
Appendix
Parameter group / operation
Stage A
Stage B
Stage C
Memory tokens Emem
train
freeze
freeze
Poolers AttnPoolq,AttnPools
train
freeze
freeze
Encoders fq,fs
train
freeze
freeze
Skill Decoder Ds
train
freeze
freeze
Token Projector Tψ
train
freeze
freeze
Backbone Fθ
last layers
train
freeze
Appendix
Table 5: Stage-wise optimization and memory-bank operations. “Last layers” denotes the partial backbone training in Stage A. The optional Skill Adapter is a Stage-B component; it is fixed or inactive in Stage C.
Category
Hyperparameter
Value
Architecture
Memory tokens m
4
Query dimension dq
256
Skill dimension ds
256
Skill tokens ns
4
Memory bank / retrieval
Initial bank size K0
128
Retrieval count Kr
2
Appendix
Table 6: Reported default hyperparameters for Optimus-R.
Task
RDT
ACT
DP
DP3
π0
π0.5†
Optimus-R
Adjust Bottle
75%
23%
0%
3%
56%
75%
89%
Click Alarmclock
12%
4%
5%
14%
11%
44%
41%
Click Bell
9%
3%
0%
0%
3%
64%
53%
Grab Roller
43%
25%
0%
2%
80%
82%
94%
Move Playingcard Away
11%
0%
0%
3%
22%
32%
38%
Pick Diverse Bottles
0%
0%
0%
1%
6%
29%
36%
Appendix
Table 7: Performance comparison on RoboTwin 2.0. We report per-task success rates (SR) over 100 rollouts under Hard setting. † represents the result we reproduced.
Cat.
Task Family
Arm Mode
Objects
Targets
A
Reorientation
Single-arm
bottle, mug
-
Block Stacking
Single-arm
block
-
Place Object onto Plate
Single-arm
can, bowl, fruit
plate
Place Object onto Tablecloth
Single-arm
bowl, plate, cup
tablecloth
B
Put Objects into Container
Bimanual
bottle, can
basket
Bimanual Basket Placement
Bimanual
basket
tablecloth
Appendix
Table 8: Detailed configuration of the real-world tasks for lifelong learning evaluation. The tasks are divided into Base Tasks (Category A) and Novel Tasks (Category B), featuring a transition from single-arm to bimanual manipulation.
Figure 7: Qualitative rollouts on real-world Category-A tasks. We visualize representative trajectories for four previously learned manipulation tasks: Reorientation , Block Stacking , Place Object onto Plate , and Place Object onto Tablecloth . Across these tasks, Optimus-R consistently executes precise object-centric manipulation behaviors under diverse object and target configurations. These results illustrate that the learned query-skill memory can preserve reusable visuomotor primitives and support stable execution of previously acquired skills during subsequent adaptation.
Figure 8: Qualitative rollouts on real-world Category-B tasks under the lifelong adaptation setting. We show representative trajectories for five newly introduced tasks: Put Objects into Container , Bimanual Basket Placement , Multi-object Collection , Sequential Pick and Place , and Bimanual Coordination Sequential Task . These tasks require multi-object reasoning, long-horizon sequencing, and coordinated dual-arm control, posing a substantial distribution shift from the Category-A task set. Optimus-R successfully adapts to these novel behaviors by retrieving, updating, and expanding explicit query-skill memory entries, enabling new skill acquisition while mitigating interference with previously learned manipulation capabilities.
Variant
Average SR (%)
Optimus-R (Full)
97.8
w/o skill-aware clustering
95.8
w/o exemplar storage
96.3
Uniform downsampling
97.2
Appendix
Table 9: Memory-bank component ablations on LIBERO with 100% training data.
K0
Kr
m
Average SR (%)
64
2
4
97.2
128
2
4
97.8
256
2
4
97.8
128
1
4
97.0
128
4
8
97.4
128
4
1
96.8
Appendix
Table 10: LIBERO configuration sensitivity. Bold entries identify the default configuration. The last two rows jointly vary retrieval count and memory-token count.
Adaptation stage
K
New
Latency (ms)
Recall@5 (%)
After LIBERO-90
81
–
1.80
95.6
+ Spatial
84
3
1.80
95.4
+ Object
90
6
1.83
96.1
+ Goal
99
9
1.86
95.2
+ Long
109
10
1.90
94.7
Appendix
Table 11: Memory growth during sequential suite adaptation. Prototype increments are relative to the preceding stage; latency measures retrieval rather than full policy inference.
Evaluation stage
Goal SR (%)
Bowl-on-plate SR (%)
K
Zero-shot
8
68
16
After Goal adaptation
96
98
27
Appendix
Table 12: Spatial-to-Goal transfer. Bowl-on-plate denotes the Goal task put the bowl on the plate .
Checkpoint
T1
T2
T3
T4
T5
T6
T7
T8
T9
T10
After T1
100
–
–
–
–
–
–
–
–
–
After T2
98
100
–
–
–
–
–
–
–
–
After T3
98
100
100
–
–
–
–
–
–
–
After T4
100
98
100
98
–
–
–
–
–
–
After T5
100
100
94
100
98
–
–
–
–
–
After T6
96
100
96
98
94
98
–
–
–
–
Appendix
Table 13: Task-wise continual learning on LIBERO-Spatial. Entries are success rates (%); dashes denote tasks not yet introduced. T1 – T10 follow the sequence in Liu et al. (2026) .