Vision-language-action policies often see only one or a few recent frames, which makes it difficult to evaluate how they use information that disappears during a task. We introduce MIKASA-Robo-VLA, a benchmark of 90 language-conditioned manipulation tasks. All but 10 hide the cue an action depends on. Those 10 are reactive controls. MIKASA-Robo, the suite it rebuilds, has 32 tasks and uses language only in a representative VLA subset. Here every task provides an instruction, while memory-dependent tasks hide a task-relevant cue and reactive controls keep it available. For 70 tasks, environment phase timings specify an information gap, and for 28 of them the gap exceeds the 16-frame window of the widest fixed-context VLA we survey. The gap counts only the interval the cue is provably absent, not the full duration a policy must retain it, so every memory-dependent task still requires memory by construction, including the ones whose measured gap is short. We release 22,500 oracle trajectories across 10 memory types in RLDS and LeRobotDataset v3. A reference π0.5 baseline with current images and proprioception, but no observation history or explicit memory module, is fine-tuned on 14 tasks and achieves 0.211 ± 0.044 mean task success. Its lower success on the evaluated Long-split tasks is confounded by open-loop chunking and the memory types represented in that subset. Project page: https://mikasarobo.github.io/
Figures & tables
Figure 1: Memory-dependent manipulation in MIKASA-Robo-VLA. Each row shows 8 evenly spaced frames from a 600-step rollout horizon, together with the task and language instruction. Ticks locate each frame at its nominal environment step. The relevant cue disappears before the dependent action begins. Images show the top-camera view at render resolution, and policy inputs have resolution 128×128×3 . Appendix K provides rollouts for all tasks.
Memory type
Definition
Tasks
Episodes
Min t
Mean t
Median t
Max t
Object
The cue reveals which object identity (color or shape) is the target. The agent must retain that identity while acting among visually similar distractors.
18
4,500
10
178
53
599
Spatial
The goal is a place rather than an object identity: which container hides the ball, where a moving ball will arrive, a target angle, or where an object began. Where that place is hidden, the agent must recall it rather than read it off the current frame. 10 of the 14 are reactive controls whose cue is never hidden, so recall is demanded only in the other 4.
14
3,500
10
30
28
90
Capacity
The cue presents a set of items to be collected or touched in any order. The agent must retain the full set membership across the interaction, with set size scaling the memory load.
12
3,000
121
431
355
1199
Temporal
The cue is a countable or timed event (e.g., a number of blinks). The agent must retain a count or duration and reproduce it during the action phase.
12
3,000
57
338
204
1004
Negative
The cue presents a set of items and the goal is defined by exclusion. The agent must retain the whole set it was shown in order to touch the one later item whose color or shape was not in it.
9
2,250
10
15
15
28
Sequential
The cue presents an ordered sequence of items. The agent must retain both the set and the order and reproduce that order during the action phase.
6
1,500
131
489
371
1199
Table 1: Memory-type taxonomy: definitions, task and episode counts, and the realized episode-length distribution in steps (minimum, mean, median, maximum). Figure 1 shows a rollout from two of these types. Appendix K shows one for every task. Each swatch is that type’s color in every figure of this paper.
Split
Tasks
Episodes
Frames
Median length (steps)
Max length (steps)
Horizon range
Short
38
9,500
303,527
18
197
25–200
Medium
30
7,500
2,154,062
272
599
250–600
Long
22
5,500
3,849,364
698
1,523
700–2,160
Table 2: Horizon-split composition: task and episode counts, frame totals, realized episode length, and the design horizon range.
Figure 2: Realized episode lengths across 22,500 demonstrations. (a) Empirical cumulative distributions for the full release and each horizon split, with a logarithmic horizontal axis. (b) Length distributions by memory type, ordered by median. Violins show kernel-density estimates. Horizontal lines indicate the observed range and ticks mark the median.
Task
Memory type
Horizon
Success
BunchOfColors3-Long-VLA-v0
Capacity
Long
0.000
ChainOfColors3-Long-VLA-v0
Sequential
Long
0.000
GatherAndRecall1-VLA-v0
Prospective
Short
0.350
GatherAndRecall3-VLA-v0
Prospective
Medium
0.150
InterceptGrabMedium-VLA-v0
Spatial
Short
0.050
InterceptMedium-VLA-v0
Spatial
Short
0.400
Table 3: Reference π0.5 performance with current images and proprioception but no observation history or explicit memory module, evaluated on 20 episodes per task. The aggregate is the unweighted mean across tasks.
Appendix figures & tables72 assets
Supplementary material from the paper’s appendix.
Appendix
Key
Type
Notes
observation.images.top
video
128×128×3 , AV1 in yuv420p , 10Hz , static camera
observation.images.wrist
video
as above, wrist-mounted camera
observation.state
float32 [7]
physical units, not normalized
action
float32 [7]
normalized to [−1,1]
timestamp
float32 [1]
seconds from episode start
frame_index
int64 [1]
step index within the episode
Appendix
Table 4: LeRobotDataset v3 feature schema, from each task’s meta/info.json . Video frames decode to HWC uint8 . The declared shape is the frame shape, not the per-row shape.
Key under steps
Type
Notes
observation/image
uint8 128×128×3
PNG-encoded, → observation.images.top
observation/wrist_image
uint8 128×128×3
PNG-encoded, → observation.images.wrist
observation/proprio
float32 [7]
→ observation.state
action
float32 [7]
normalized to [−1,1]
reward
float32
no LeRobot counterpart, see below
discount
float32
written as the constant 1.0 on every step, so it carries nothing
Appendix
Table 5: RLDS per-step schema, from each task’s features.json . The Notes column gives the corresponding LeRobot key where one exists.
Task
k
1/k
Success
GatherAndRecall1-VLA-v0
3
0.333
0.35
GatherAndRecall3-VLA-v0
3
0.333
0.15
RememberColor5-Long-VLA-v0
5
0.200
0.30
RememberColor5-VLA-v0
5
0.200
0.25
RememberShape5-VLA-v0
5
0.200
0.05
RememberShapeAndColor3x2-VLA-v0
6
0.167
0.20
Appendix
Table 6: Selection floors for the 7 evaluated tasks whose success condition names a finite candidate set. k is the number of candidates the success condition selects among, never transcribed by hand. Success is the reference arm’s rate over 20 episodes. A rate below 1/k does not mean the arm is worse than forgetful: at this episode count the two are not separable. Nor does a rate above it mean the arm is better – 0 of these rows clear their floor by more than the row’s own standard error.
Figure 3: Rollout filmstrip, BatteriesCheckerEasy-3-VLA-v0 (a representative variant).
Env ID
Horizon (steps)
Source
Median length
Gap
BatteriesCheckerEasy-3-VLA-v0
540
MP
363
undetermined
BatteriesCheckerEasy-6-VLA-v0
1080
MP
727
undetermined
Appendix
Table 7: Batteries Checker Easy: variants, horizon, data source, median realized episode length, and information gap. Lengths and gaps are in environment steps. Every variant contributes 250 retained demonstrations, and their oracle success rate is 1.000. Each horizon’s split is given by Table 2 .
Figure 4: Rollout filmstrip, BatteriesCheckerHard-3-VLA-v0 (a representative variant).
Env ID
Horizon (steps)
Source
Median length
Gap
BatteriesCheckerHard-3-VLA-v0
1080
MP
699
undetermined
BatteriesCheckerHard-6-VLA-v0
2160
MP
1416
undetermined
Appendix
Table 8: Batteries Checker Hard: variants, horizon, data source, median realized episode length, and information gap. Lengths and gaps are in environment steps. Every variant contributes 250 retained demonstrations, and their oracle success rate is 1.000. Each horizon’s split is given by Table 2 .
Figure 5: Rollout filmstrip, BlinkCountButtonPressEasy-VLA-v0 (a representative variant).
Env ID
Horizon (steps)
Source
Median length
Gap
BlinkCountButtonPressEasy-VLA-v0
150
MP
94
0
BlinkCountButtonPressMedium-VLA-v0
200
MP
121
0
BlinkCountButtonPressHard-VLA-v0
300
MP
153
0
BlinkCountButtonPressEasy-Long-VLA-v0
1200
MP
199
0
BlinkCountButtonPressMedium-Long-VLA-v0
1200
MP
490
0
BlinkCountButtonPressHard-Long-VLA-v0
1200
MP
788
0
Appendix
Table 9: Blink Count Button Press: variants, horizon, data source, median realized episode length, and information gap. Lengths and gaps are in environment steps. Every variant contributes 250 retained demonstrations, and their oracle success rate is 1.000. Each horizon’s split is given by Table 2 .
Figure 6: Rollout filmstrip, BunchOfColors3-VLA-v0 (a representative variant).
Env ID
Horizon (steps)
Source
Median length
Gap
BunchOfColors3-VLA-v0
400
MP
158
3 (2–4)
BunchOfColors5-VLA-v0
400
MP
242
3 (2–4)
BunchOfColors7-VLA-v0
400
MP
327
3 (2–4)
BunchOfColors3-Long-VLA-v0
700
MP
449
225 (138–313)
BunchOfColors5-Long-VLA-v0
700
MP
524
225 (138–313)
BunchOfColors7-Long-VLA-v0
700
MP
569
225 (138–313)
Appendix
Table 10: Bunch Of Colors: variants, horizon, data source, median realized episode length, and information gap. Lengths and gaps are in environment steps. Every variant contributes 250 retained demonstrations, and their oracle success rate is 1.000. Each horizon’s split is given by Table 2 .
Figure 7: Rollout filmstrip, ChainOfColors3-VLA-v0 (a representative variant).
Env ID
Horizon (steps)
Source
Median length
Gap
ChainOfColors3-VLA-v0
400
MP
165
3 (2–4)
ChainOfColors5-VLA-v0
400
MP
253
3 (2–4)
ChainOfColors7-VLA-v0
400
MP
342
3 (2–4)
ChainOfColors3-Long-VLA-v0
800
MP
559
225 (138–313)
ChainOfColors5-Long-VLA-v0
1000
MP
740
225 (138–313)
ChainOfColors7-Long-VLA-v0
1200
MP
916
225 (138–313)
Appendix
Table 11: Chain Of Colors: variants, horizon, data source, median realized episode length, and information gap. Lengths and gaps are in environment steps. Every variant contributes 250 retained demonstrations, and their oracle success rate is 1.000. Each horizon’s split is given by Table 2 .
Figure 8: Rollout filmstrip, FindImposterColor3-VLA-v0 (a representative variant).
Env ID
Horizon (steps)
Source
Median length
Gap
FindImposterColor3-VLA-v0
25
PPO
15
3 (2–4)
FindImposterColor5-VLA-v0
25
PPO
15
3 (2–4)
FindImposterColor9-VLA-v0
25
PPO
16
3 (2–4)
Appendix
Table 12: Find Imposter Color: variants, horizon, data source, median realized episode length, and information gap. Lengths and gaps are in environment steps. Every variant contributes 250 retained demonstrations, and their oracle success rate is 1.000. Each horizon’s split is given by Table 2 .
Figure 9: Rollout filmstrip, FindImposterShape3-VLA-v0 (a representative variant).
Env ID
Horizon (steps)
Source
Median length
Gap
FindImposterShape3-VLA-v0
25
PPO
14
3 (2–4)
FindImposterShape5-VLA-v0
25
PPO
15
3 (2–4)
FindImposterShape9-VLA-v0
25
PPO
15
3 (2–4)
Appendix
Table 13: Find Imposter Shape: variants, horizon, data source, median realized episode length, and information gap. Lengths and gaps are in environment steps. Every variant contributes 250 retained demonstrations, and their oracle success rate is 1.000. Each horizon’s split is given by Table 2 .
Figure 10: Rollout filmstrip, FindImposterShapeAndColor3x2-VLA-v0 (a representative variant).
Env ID
Horizon (steps)
Source
Median length
Gap
FindImposterShapeAndColor3x2-VLA-v0
25
PPO
16
3 (2–4)
FindImposterShapeAndColor3x3-VLA-v0
25
PPO
15
3 (2–4)
FindImposterShapeAndColor5x3-VLA-v0
40
PPO
15
3 (2–4)
Appendix
Table 14: Find Imposter Shape And Color: variants, horizon, data source, median realized episode length, and information gap. Lengths and gaps are in environment steps. Every variant contributes 250 retained demonstrations, and their oracle success rate is 1.000. Each horizon’s split is given by Table 2 .
Figure 11: Rollout filmstrip, GatherAndRecall1-VLA-v0 (a representative variant).
Env ID
Horizon (steps)
Source
Median length
Gap
GatherAndRecall1-VLA-v0
200
MP
115
undetermined
GatherAndRecall3-VLA-v0
400
MP
318
undetermined
GatherAndRecall5-VLA-v0
600
MP
514
undetermined
GatherAndRecall7-VLA-v0
800
MP
721
undetermined
GatherAndRecall9-VLA-v0
1000
MP
900
undetermined
Appendix
Table 15: Gather And Recall: variants, horizon, data source, median realized episode length, and information gap. Lengths and gaps are in environment steps. Every variant contributes 250 retained demonstrations, and their oracle success rate is 1.000. Each horizon’s split is given by Table 2 .
Figure 12: Rollout filmstrip, InterceptSlow-VLA-v0 (a representative variant).
Env ID
Horizon (steps)
Source
Median length
Gap
InterceptSlow-VLA-v0
60
PPO
43
cue never hidden
InterceptMedium-VLA-v0
60
PPO
41
cue never hidden
InterceptFast-VLA-v0
60
PPO
30
cue never hidden
Appendix
Table 16: Intercept: variants, horizon, data source, median realized episode length, and information gap. Lengths and gaps are in environment steps. Every variant contributes 250 retained demonstrations, and their oracle success rate is 1.000. Each horizon’s split is given by Table 2 .
Figure 13: Rollout filmstrip, InterceptGrabSlow-VLA-v0 (a representative variant).
Env ID
Horizon (steps)
Source
Median length
Gap
InterceptGrabSlow-VLA-v0
60
PPO
24
cue never hidden
InterceptGrabMedium-VLA-v0
60
PPO
33
cue never hidden
InterceptGrabFast-VLA-v0
60
PPO
50
cue never hidden
Appendix
Table 17: Intercept Grab: variants, horizon, data source, median realized episode length, and information gap. Lengths and gaps are in environment steps. Every variant contributes 250 retained demonstrations, and their oracle success rate is 1.000. Each horizon’s split is given by Table 2 .
Figure 14: Rollout filmstrip, RememberColor3-VLA-v0 (a representative variant).
Env ID
Horizon (steps)
Source
Median length
Gap
RememberColor3-VLA-v0
25
PPO
15
3 (2–4)
RememberColor5-VLA-v0
25
PPO
15
3 (2–4)
RememberColor9-VLA-v0
25
PPO
15
3 (2–4)
RememberColor3-Long-VLA-v0
600
MP
337
250 (150–350)
RememberColor5-Long-VLA-v0
600
MP
390
250 (150–350)
RememberColor9-Long-VLA-v0
600
MP
378
250 (150–350)
Appendix
Table 18: Remember Color: variants, horizon, data source, median realized episode length, and information gap. Lengths and gaps are in environment steps. Every variant contributes 250 retained demonstrations, and their oracle success rate is 1.000. Each horizon’s split is given by Table 2 .
Figure 15: Rollout filmstrip, RememberShape3-VLA-v0 (a representative variant).
Env ID
Horizon (steps)
Source
Median length
Gap
RememberShape3-VLA-v0
25
PPO
15
3 (2–4)
RememberShape5-VLA-v0
25
PPO
15
3 (2–4)
RememberShape9-VLA-v0
25
PPO
15
3 (2–4)
RememberShape3-Long-VLA-v0
600
MP
337
250 (150–350)
RememberShape5-Long-VLA-v0
600
MP
349
250 (150–350)
RememberShape9-Long-VLA-v0
600
MP
328
250 (150–350)
Appendix
Table 19: Remember Shape: variants, horizon, data source, median realized episode length, and information gap. Lengths and gaps are in environment steps. Every variant contributes 250 retained demonstrations, and their oracle success rate is 1.000. Each horizon’s split is given by Table 2 .
Figure 16: Rollout filmstrip, RememberShapeAndColor3x2-VLA-v0 (a representative variant).
Env ID
Horizon (steps)
Source
Median length
Gap
RememberShapeAndColor3x2-VLA-v0
25
PPO
15
3 (2–4)
RememberShapeAndColor3x3-VLA-v0
25
PPO
14
3 (2–4)
RememberShapeAndColor5x3-VLA-v0
25
PPO
15
3 (2–4)
RememberShapeAndColor3x2-Long-VLA-v0
600
MP
312
250 (150–350)
RememberShapeAndColor3x3-Long-VLA-v0
600
MP
321
250 (150–350)
RememberShapeAndColor5x3-Long-VLA-v0
600
MP
322
250 (150–350)
Appendix
Table 20: Remember Shape And Color: variants, horizon, data source, median realized episode length, and information gap. Lengths and gaps are in environment steps. Every variant contributes 250 retained demonstrations, and their oracle success rate is 1.000. Each horizon’s split is given by Table 2 .
Figure 17: Rollout filmstrip, RotateLenientPos-VLA-v0 (a representative variant).
Env ID
Horizon (steps)
Source
Median length
Gap
RotateLenientPos-VLA-v0
60
PPO
29
cue never hidden
RotateLenientPosNeg-VLA-v0
60
PPO
19
cue never hidden
Appendix
Table 21: Rotate Lenient: variants, horizon, data source, median realized episode length, and information gap. Lengths and gaps are in environment steps. Every variant contributes 250 retained demonstrations, and their oracle success rate is 1.000. Each horizon’s split is given by Table 2 .
Figure 18: Rollout filmstrip, RotateStrictPos-VLA-v0 (a representative variant).
Env ID
Horizon (steps)
Source
Median length
Gap
RotateStrictPos-VLA-v0
90
PPO
31
cue never hidden
RotateStrictPosNeg-VLA-v0
90
PPO
22
cue never hidden
Appendix
Table 22: Rotate Strict: variants, horizon, data source, median realized episode length, and information gap. Lengths and gaps are in environment steps. Every variant contributes 250 retained demonstrations, and their oracle success rate is 1.000. Each horizon’s split is given by Table 2 .
Figure 19: Rollout filmstrip, SeqOfColors3-VLA-v0 (a representative variant).
Env ID
Horizon (steps)
Source
Median length
Gap
SeqOfColors3-VLA-v0
400
MP
165
3 (2–4)
SeqOfColors5-VLA-v0
400
MP
253
3 (2–4)
SeqOfColors7-VLA-v0
400
MP
342
3 (2–4)
SeqOfColors3-Long-VLA-v0
800
MP
559
225 (138–313)
SeqOfColors5-Long-VLA-v0
1000
MP
738
225 (138–313)
SeqOfColors7-Long-VLA-v0
1200
MP
916
225 (138–313)
Appendix
Table 23: Seq Of Colors: variants, horizon, data source, median realized episode length, and information gap. Lengths and gaps are in environment steps. Every variant contributes 250 retained demonstrations, and their oracle success rate is 1.000. Each horizon’s split is given by Table 2 .
Figure 20: Rollout filmstrip, ShellGameColorLampTouch-VLA-v0 (a representative variant).
Env ID
Horizon (steps)
Source
Median length
Gap
ShellGameColorLampTouch-VLA-v0
30
PPO
15
0
Appendix
Table 24: Shell Game Color Lamp Touch: variants, horizon, data source, median realized episode length, and information gap. Lengths and gaps are in environment steps. Every variant contributes 250 retained demonstrations, and their oracle success rate is 1.000. Each horizon’s split is given by Table 2 .
Figure 21: Rollout filmstrip, ShellGamePush-VLA-v0 (a representative variant).
Env ID
Horizon (steps)
Source
Median length
Gap
ShellGamePush-VLA-v0
30
PPO
15
0
Appendix
Table 25: Shell Game Push: variants, horizon, data source, median realized episode length, and information gap. Lengths and gaps are in environment steps. Every variant contributes 250 retained demonstrations, and their oracle success rate is 1.000. Each horizon’s split is given by Table 2 .
Figure 22: Rollout filmstrip, ShellGameShuffleColorLampTouch-VLA-v0 (a representative variant).
Env ID
Horizon (steps)
Source
Median length
Gap
ShellGameShuffleColorLampTouch-VLA-v0
60
PPO
44
28 (24–31)
ShellGameShuffleColorLampTouch-Long-VLA-v0
600
MP
331
250 (175–325)
Appendix
Table 26: Shell Game Shuffle Color Lamp Touch: variants, horizon, data source, median realized episode length, and information gap. Lengths and gaps are in environment steps. Every variant contributes 250 retained demonstrations, and their oracle success rate is 1.000. Each horizon’s split is given by Table 2 .
Figure 23: Rollout filmstrip, ShellGameShuffleTouch-VLA-v0 (a representative variant).
Env ID
Horizon (steps)
Source
Median length
Gap
ShellGameShuffleTouch-VLA-v0
60
PPO
42
28 (24–31)
ShellGameShuffleTouch-Long-VLA-v0
600
MP
337
250 (175–325)
Appendix
Table 27: Shell Game Shuffle Touch: variants, horizon, data source, median realized episode length, and information gap. Lengths and gaps are in environment steps. Every variant contributes 250 retained demonstrations, and their oracle success rate is 1.000. Each horizon’s split is given by Table 2 .
Figure 24: Rollout filmstrip, ShellGameTouch-VLA-v0 (a representative variant).
Env ID
Horizon (steps)
Source
Median length
Gap
ShellGameTouch-VLA-v0
30
PPO
22
0
Appendix
Table 28: Shell Game Touch: variants, horizon, data source, median realized episode length, and information gap. Lengths and gaps are in environment steps. Every variant contributes 250 retained demonstrations, and their oracle success rate is 1.000. Each horizon’s split is given by Table 2 .
Figure 25: Rollout filmstrip, TakeItBack-VLA-v0 (a representative variant).
Env ID
Horizon (steps)
Source
Median length
Gap
TakeItBack-VLA-v0
60
PPO
27
undetermined
Appendix
Table 29: Take It Back: variants, horizon, data source, median realized episode length, and information gap. Lengths and gaps are in environment steps. Every variant contributes 250 retained demonstrations, and their oracle success rate is 1.000. Each horizon’s split is given by Table 2 .
Figure 26: Rollout filmstrip, TimedTransferEasy-VLA-v0 (a representative variant).
Env ID
Horizon (steps)
Source
Median length
Gap
TimedTransferEasy-VLA-v0
200
MP
106
100
TimedTransferMedium-VLA-v0
250
MP
152
150
TimedTransferHard-VLA-v0
300
MP
203
200
TimedTransferEasy-Long-VLA-v0
600
MP
298
300
TimedTransferMedium-Long-VLA-v0
900
MP
488
500
TimedTransferHard-Long-VLA-v0
1200
MP
962
1000
Appendix
Table 30: Timed Transfer: variants, horizon, data source, median realized episode length, and information gap. Lengths and gaps are in environment steps. Every variant contributes 250 retained demonstrations, and their oracle success rate is 1.000. Each horizon’s split is given by Table 2 .
Figure 27: Rollout filmstrip, TraceShapeEasy-VLA-v0 (a representative variant).
Env ID
Horizon (steps)
Source
Median length
Gap
TraceShapeEasy-VLA-v0
250
MP
209
0
TraceShapeMedium-VLA-v0
300
MP
210
0
TraceShapeHard-VLA-v0
350
MP
207
0
Appendix
Table 31: Trace Shape: variants, horizon, data source, median realized episode length, and information gap. Lengths and gaps are in environment steps. Every variant contributes 250 retained demonstrations, and their oracle success rate is 1.000. Each horizon’s split is given by Table 2 .
Figure 28: Rollout filmstrip, TraceShapeSeqEasy-VLA-v0 (a representative variant).
Env ID
Horizon (steps)
Source
Median length
Gap
TraceShapeSeqEasy-VLA-v0
1500
MP
702
0
TraceShapeSeqMedium-VLA-v0
1500
MP
720
0
TraceShapeSeqHard-VLA-v0
1500
MP
706
0
Appendix
Table 32: Trace Shape Seq: variants, horizon, data source, median realized episode length, and information gap. Lengths and gaps are in environment steps. Every variant contributes 250 retained demonstrations, and their oracle success rate is 1.000. Each horizon’s split is given by Table 2 .
Figure 29: Total released frames per task, sorted descending, colored by memory type. The same color maps to the same memory type in every figure in this paper (Table 1 ). Numeric values for every task are in Table 33 below.
Env ID
Memory type
Split
Horizon
Source
Mean len.
Median len.
Oracle succ.
Gap
ShellGameTouch-VLA-v0
Spatial
Short
30
PPO
21.5
22
1.000
0
ShellGamePush-VLA-v0
Spatial
Short
30
PPO
15.0
15
1.000
0
InterceptSlow-VLA-v0
Spatial
Short
60
PPO
43.8
43
1.000
cue never hidden
InterceptMedium-VLA-v0
Spatial
Short
60
PPO
41.8
41
1.000
cue never hidden
InterceptFast-VLA-v0
Spatial
Short
60
PPO
31.6
30
1.000
cue never hidden
InterceptGrabSlow-VLA-v0
Spatial
Short
60
PPO
27.6
24
1.000
cue never hidden
Appendix
Table 33: Full per-task statistics for all 90 canonical tasks. Every task contributes 250 episodes. Lengths, horizons, and gaps are in environment steps. In the Gap column a bare number is a constant the environment sets, median (p25–p75) is that constant drawn afresh each episode, undetermined marks an interval that ends when the agent reaches some state rather than after a stated number of steps, and cue never hidden marks a task whose source states no interval exists.
openpi-comet [ Bai et al., 2025 ] , a JAX/Flax fork of openpi
Training data
LeRobot format, 14 tasks × 250 episodes
(3,500 episodes, 798,296 frames)
Image input
Two cameras stacked, 128×128×3 each, upscaled to 224×224
Action space
pd_ee_delta_pose , 7 active dims, padded to 32
Appendix
Table 34: Reference π0.5 baseline training configuration.
Scale
Controlled memory difficulty
Environment & learning
Benchmark
Backend
Tasks
Types
Demos
Ret.
Load
H-strat.
Dense rew.
GPU vec.
MemoryBench [ Fang et al., 2025 ]
RLBench / Coppelia
3
2
300
✗
✗
✗
✗
✗
LIBERO-Mem [ Chung et al., 2026 ]
LIBERO / MuJoCo
10
4
1,000
∘
∘
∘
✗
✗
RMBench [ Chen et al., 2026 ]
RoboTwin / SAPIEN
9
–
450
✗
∘
✗
✗
✗
RoboMME [ Dai et al., 2026 ]
ManiSkill / SAPIEN
16
4
1,600
✗
✗
✗
✗
✗
RoboMemArena [ Lei et al., 2026 ]
LIBERO / MuJoCo
26
4
2,600
✗
✗
✗
✗
✗
Appendix
Table 35: Comparison of dedicated simulation benchmarks for memory-dependent robotic manipulation.
Task
No memory
μ VLA
InterceptFast
0.00
0.29
InterceptGrabFast
0.00
0.00
InterceptGrabMedium
0.00
0.00
InterceptGrabSlow
0.00
0.00
InterceptMedium †
0.36
0.55
InterceptSlow
0.05
0.08
Appendix
Table 36: Per-task success on the 23-task evaluation reported by Cherepanov et al. [2026b] : OpenVLA-OFT without a memory module (exec-8) versus μ VLA, the same backbone with a 64-token recurrent memory (receding horizon, memory updated every step). 100 episodes per task, checkpoint step 150000. Trained tasks are marked † . The other 18 are zero-shot/held-out for both arms. Evaluated under an earlier release protocol (see Section 6 ). n/a marks an evaluation that did not complete and so is not scored.
Loader
Regime
ShellGamePush
InterceptMedium
TakeItBack
RememberColor5
RSC3x3
Mean
A (shuffled)
exec-8
0.93
0.30
0.68
0.14
0.07
0.424
A (shuffled)
exec-1
0.28
0.42
0.77
0.28
0.14
0.378
B (episodic)
exec-8
0.95
0.31
0.77
0.18
0.03
0.448
B (episodic)
exec-1
0.33
0.51
0.81
0.18
0.15
0.396
Appendix
Table 37: Reference π0.5 (LeRobot port), two independently trained checkpoints, each evaluated under two execution regimes at inference: exec-8 (open-loop chunk of 8 actions) and exec-1 (receding horizon, one action per inference). 100 episodes per task per cell. Not part of the canonical 90-task benchmark evaluation.
Task
K=2 (16 steps, step 140,000)
K=8 (64 steps, best step 42,500)
BunchOfColors3
0.00
0.00
ChainOfColors3
0.00
0.00
GatherAndRecall3
0.24
0.26
RememberColor3-Medium
0.32
0.52
RememberColor5-Medium
0.08
0.54
RememberShapeAndColor3x2-Medium
0.18
0.44
Appendix
Table 38: Benchmark-customization case study: μ VLA with memory updated once per 8-action chunk (full open loop), trained on 3 canonical tasks plus 5 custom *-Medium-VLA-v0 variants created for this test and excluded from every count of the canonical 90-task suite. K is the TBPTT window in chunks (env-steps =8K ). 50 episodes per task.
Figure 30: Rollout filmstrips for the 18 Object tasks, one row per task, 8 evenly spaced frames per rollout, left to right. The header color is this memory type’s color throughout the paper (Table 1 ).
Figure 31: Rollout filmstrips for the 14 Spatial tasks, one row per task, 8 evenly spaced frames per rollout, left to right. The header color is this memory type’s color throughout the paper (Table 1 ).
Figure 32: Rollout filmstrips for the 12 Capacity tasks, one row per task, 8 evenly spaced frames per rollout, left to right. The header color is this memory type’s color throughout the paper (Table 1 ).
Figure 33: Rollout filmstrips for the 12 Temporal tasks, one row per task, 8 evenly spaced frames per rollout, left to right. The header color is this memory type’s color throughout the paper (Table 1 ).
Figure 34: Rollout filmstrips for the 9 Negative tasks, one row per task, 8 evenly spaced frames per rollout, left to right. The header color is this memory type’s color throughout the paper (Table 1 ).
Figure 35: Rollout filmstrips for the 6 Sequential tasks, one row per task, 8 evenly spaced frames per rollout, left to right. The header color is this memory type’s color throughout the paper (Table 1 ).
Figure 36: Rollout filmstrips for the 6 Procedural tasks, one row per task, 8 evenly spaced frames per rollout, left to right. The header color is this memory type’s color throughout the paper (Table 1 ).
Figure 37: Rollout filmstrips for the 5 Prospective tasks, one row per task, 8 evenly spaced frames per rollout, left to right. The header color is this memory type’s color throughout the paper (Table 1 ).
Figure 38: Rollout filmstrips for the 4 Tracking tasks, one row per task, 8 evenly spaced frames per rollout, left to right. The header color is this memory type’s color throughout the paper (Table 1 ).
Figure 39: Rollout filmstrips for the 4 Checklist tasks, one row per task, 8 evenly spaced frames per rollout, left to right. The header color is this memory type’s color throughout the paper (Table 1 ).
Vision-language-action (VLA) models have driven rapid progress in robotic manipulation, demonstrating strong fine-grained control and promising performance on long-horizon tasks. However, many existing VLAs lack explicit access to interaction history, making them vulnerable to perceptual aliasing: similar current observations and robot states at different task stages may induce action ambiguity and lower success rate. Existing methods incorporate temporal or progress cues through feature conditioning, action-prior modification, or sampling guidance. However, methods that jointly fine-tune memory modules and the base VLA incur additional policy-training costs, motivating the separation of trainable history-conditioned steering from frozen base-policy refinement. We propose ActMem-VLA, a dual-expert handover architecture that augments a frozen, fine-tuned VLA with a memory plugin comprising a Mamba-based memory module and a lightweight PreAction Expert (PAE). Specifically, Mamba encodes executed-action history into memory that conditions PAE alongside current context. With these inputs, PAE steers task progression during early, high-noise denoising, then passes the partially denoised action to the frozen Action Expert (AE) to refine action details during the remaining low-noise steps. The fine-tuned base VLA remains frozen throughout training, while only the Mamba module and PAE are jointly optimized. On LIBERO-Mem, ActMem-VLA achieves 80.8% average success across all ten tasks, compared with 65.2% for π0.5 and 49.5% for MemoryVLA, while introducing only 3.45% additional parameters. Across four real-world tasks, it improves the average success rate over π0.5 by 28.8%.
Yaxin Zhao, Dianye Huang, Chenwei Wang +2
Harbin Institute of Technology, Harbin, China. · Medical Intelligence and Robotic Cognition (MIRoC) Lab, Department of Mechanical Engineering, The University of Hong Kong (HKU), Hong Kong SAR, China. · Department of Computing, The Hong Kong Polytechnic University, Hong Kong SAR, China
Memory is essential for long-horizon, partially observed robotic manipulation: a robot must remember which object was placed in a drawer, whose cup it moved, or how many action cycles have elapsed. Recent vision-language-action (VLA) models embed memory directly inside the policy, but benchmarks show no single in-policy mechanism covers all spatio-temporal dimensions, trailing oracle methods by a wide margin. We argue that memory type dictates where memory should reside: short-term perceptual memory (repetition, timing, retracing) belongs inside the policy, while long-term object memory (persistent spatial state, containment, event history) belongs outside as an explicit, readable record. We present Ledger, a harness that realizes this split over a single fine-tuned π0.5 policy by pairing an in-policy frame-sampling memory with an external spatio-temporal object memory, the ledger, built from a SAM3 tracker and a VLM captioner of the demonstration and read by an LLM planner that decides at step boundaries. On RoboMME, Ledger reaches the highest four-suite average among the evaluated methods, 64.3% (vs. 45.9% for the strongest prior method under identical evaluation), leading object reference (60.7% vs. 40.3%) and object permanence (86.7% vs. 56.2%) using a single set of weights. Choosing the memory source at runtime, from the instruction and the record, removes the need for a task-level router.
Long-horizon manipulation requires robots to remember cues that are no longer in view while responding to moving objects. Yet vision-language-action (VLA) policies often rely on the latest observation, and refreshing their visual context typically requires another costly vision-language model (VLM) pass. We present D2-VLA, which combines dual memory and dual-frequency control at the KV-cache interface of a pretrained VLA. D2-VLA uses block-wise causal KV caching to encode observations incrementally and, guided by distinct temporal attention patterns, constructs separate historical KV read views for the VLM and action expert. Between periodic VLM updates, a gated adapter incorporates fresh visual features into the latest history-conditioned KV block, while a short fast-memory queue supports action replanning. We introduce DOMINO-Long, a ten-task benchmark requiring robots to use earlier visual cues when manipulating moving objects. D2-VLA achieves complete-task success rates of 29.3% on DOMINO, compared with 9.6% for π0.5 and 17.2% for PUMA, and 60.0% on DOMINO-Long, compared with 35.4% and 20.6%, respectively. It improves success rates on eight real-robot tasks and reaches 97.5% on LIBERO-Long and 74.3% on RoboTwin 2.0.
Zijian Ye, Chengqi Wei, Wei Huang +9
The University of Hong Kong · Southern University of Science and Technology