MIKASA-Robo-VLA: Benchmarking Memory in VLA Models for Long-Horizon Manipulation
Organizations: AXXX, Moscow, Russia · MIRIAI, Moscow, Russia · HSE University, Moscow, Russia
Abstract
Vision-language-action policies often see only one or a few recent frames, which makes it difficult to evaluate how they use information that disappears during a task. We introduce MIKASA-Robo-VLA, a benchmark of 90 language-conditioned manipulation tasks. All but 10 hide the cue an action depends on. Those 10 are reactive controls. MIKASA-Robo, the suite it rebuilds, has 32 tasks and uses language only in a representative VLA subset. Here every task provides an instruction, while memory-dependent tasks hide a task-relevant cue and reactive controls keep it available. For 70 tasks, environment phase timings specify an information gap, and for 28 of them the gap exceeds the 16-frame window of the widest fixed-context VLA we survey. The gap counts only the interval the cue is provably absent, not the full duration a policy must retain it, so every memory-dependent task still requires memory by construction, including the ones whose measured gap is short. We release 22,500 oracle trajectories across 10 memory types in RLDS and LeRobotDataset v3. A reference baseline with current images and proprioception, but no observation history or explicit memory module, is fine-tuned on 14 tasks and achieves 0.211 0.044 mean task success. Its lower success on the evaluated Long-split tasks is confounded by open-loop chunking and the memory types represented in that subset. Project page: https://mikasarobo.github.io/
Figures & tables
| Memory type | Definition | Tasks | Episodes | Min t | Mean t | Median t | Max t |
|---|---|---|---|---|---|---|---|
| Object | The cue reveals which object identity (color or shape) is the target. The agent must retain that identity while acting among visually similar distractors. | 18 | 4,500 | 10 | 178 | 53 | 599 |
| Spatial | The goal is a place rather than an object identity: which container hides the ball, where a moving ball will arrive, a target angle, or where an object began. Where that place is hidden, the agent must recall it rather than read it off the current frame. 10 of the 14 are reactive controls whose cue is never hidden, so recall is demanded only in the other 4. | 14 | 3,500 | 10 | 30 | 28 | 90 |
| Capacity | The cue presents a set of items to be collected or touched in any order. The agent must retain the full set membership across the interaction, with set size scaling the memory load. | 12 | 3,000 | 121 | 431 | 355 | 1199 |
| Temporal | The cue is a countable or timed event (e.g., a number of blinks). The agent must retain a count or duration and reproduce it during the action phase. | 12 | 3,000 | 57 | 338 | 204 | 1004 |
| Negative | The cue presents a set of items and the goal is defined by exclusion. The agent must retain the whole set it was shown in order to touch the one later item whose color or shape was not in it. | 9 | 2,250 | 10 | 15 | 15 | 28 |
| Sequential | The cue presents an ordered sequence of items. The agent must retain both the set and the order and reproduce that order during the action phase. | 6 | 1,500 | 131 | 489 | 371 | 1199 |
| Split | Tasks | Episodes | Frames | Median length (steps) | Max length (steps) | Horizon range |
|---|---|---|---|---|---|---|
| Short | 38 | 9,500 | 303,527 | 18 | 197 | 25–200 |
| Medium | 30 | 7,500 | 2,154,062 | 272 | 599 | 250–600 |
| Long | 22 | 5,500 | 3,849,364 | 698 | 1,523 | 700–2,160 |
| Task | Memory type | Horizon | Success |
|---|---|---|---|
| BunchOfColors3-Long-VLA-v0 | Capacity | Long | 0.000 |
| ChainOfColors3-Long-VLA-v0 | Sequential | Long | 0.000 |
| GatherAndRecall1-VLA-v0 | Prospective | Short | 0.350 |
| GatherAndRecall3-VLA-v0 | Prospective | Medium | 0.150 |
| InterceptGrabMedium-VLA-v0 | Spatial | Short | 0.050 |
| InterceptMedium-VLA-v0 | Spatial | Short | 0.400 |
Appendix figures & tables72 assets
Supplementary material from the paper’s appendix.
Appendix
| Key | Type | Notes |
|---|---|---|
| observation.images.top | video | , AV1 in yuv420p , , static camera |
| observation.images.wrist | video | as above, wrist-mounted camera |
| observation.state | float32 [7] | physical units, not normalized |
| action | float32 [7] | normalized to |
| timestamp | float32 [1] | seconds from episode start |
| frame_index | int64 [1] | step index within the episode |
| Key under steps | Type | Notes |
|---|---|---|
| observation/image | uint8 | PNG-encoded, observation.images.top |
| observation/wrist_image | uint8 | PNG-encoded, observation.images.wrist |
| observation/proprio | float32 [7] | observation.state |
| action | float32 [7] | normalized to |
| reward | float32 | no LeRobot counterpart, see below |
| discount | float32 | written as the constant 1.0 on every step, so it carries nothing |
| Task | Success | ||
|---|---|---|---|
| GatherAndRecall1-VLA-v0 | 3 | 0.333 | 0.35 |
| GatherAndRecall3-VLA-v0 | 3 | 0.333 | 0.15 |
| RememberColor5-Long-VLA-v0 | 5 | 0.200 | 0.30 |
| RememberColor5-VLA-v0 | 5 | 0.200 | 0.25 |
| RememberShape5-VLA-v0 | 5 | 0.200 | 0.05 |
| RememberShapeAndColor3x2-VLA-v0 | 6 | 0.167 | 0.20 |
| Env ID | Horizon (steps) | Source | Median length | Gap |
|---|---|---|---|---|
| BatteriesCheckerEasy-3-VLA-v0 | 540 | MP | 363 | undetermined |
| BatteriesCheckerEasy-6-VLA-v0 | 1080 | MP | 727 | undetermined |
| Env ID | Horizon (steps) | Source | Median length | Gap |
|---|---|---|---|---|
| BatteriesCheckerHard-3-VLA-v0 | 1080 | MP | 699 | undetermined |
| BatteriesCheckerHard-6-VLA-v0 | 2160 | MP | 1416 | undetermined |
| Env ID | Horizon (steps) | Source | Median length | Gap |
|---|---|---|---|---|
| BlinkCountButtonPressEasy-VLA-v0 | 150 | MP | 94 | 0 |
| BlinkCountButtonPressMedium-VLA-v0 | 200 | MP | 121 | 0 |
| BlinkCountButtonPressHard-VLA-v0 | 300 | MP | 153 | 0 |
| BlinkCountButtonPressEasy-Long-VLA-v0 | 1200 | MP | 199 | 0 |
| BlinkCountButtonPressMedium-Long-VLA-v0 | 1200 | MP | 490 | 0 |
| BlinkCountButtonPressHard-Long-VLA-v0 | 1200 | MP | 788 | 0 |
| Env ID | Horizon (steps) | Source | Median length | Gap |
|---|---|---|---|---|
| BunchOfColors3-VLA-v0 | 400 | MP | 158 | 3 (2–4) |
| BunchOfColors5-VLA-v0 | 400 | MP | 242 | 3 (2–4) |
| BunchOfColors7-VLA-v0 | 400 | MP | 327 | 3 (2–4) |
| BunchOfColors3-Long-VLA-v0 | 700 | MP | 449 | 225 (138–313) |
| BunchOfColors5-Long-VLA-v0 | 700 | MP | 524 | 225 (138–313) |
| BunchOfColors7-Long-VLA-v0 | 700 | MP | 569 | 225 (138–313) |
| Env ID | Horizon (steps) | Source | Median length | Gap |
|---|---|---|---|---|
| ChainOfColors3-VLA-v0 | 400 | MP | 165 | 3 (2–4) |
| ChainOfColors5-VLA-v0 | 400 | MP | 253 | 3 (2–4) |
| ChainOfColors7-VLA-v0 | 400 | MP | 342 | 3 (2–4) |
| ChainOfColors3-Long-VLA-v0 | 800 | MP | 559 | 225 (138–313) |
| ChainOfColors5-Long-VLA-v0 | 1000 | MP | 740 | 225 (138–313) |
| ChainOfColors7-Long-VLA-v0 | 1200 | MP | 916 | 225 (138–313) |
| Env ID | Horizon (steps) | Source | Median length | Gap |
|---|---|---|---|---|
| FindImposterColor3-VLA-v0 | 25 | PPO | 15 | 3 (2–4) |
| FindImposterColor5-VLA-v0 | 25 | PPO | 15 | 3 (2–4) |
| FindImposterColor9-VLA-v0 | 25 | PPO | 16 | 3 (2–4) |
| Env ID | Horizon (steps) | Source | Median length | Gap |
|---|---|---|---|---|
| FindImposterShape3-VLA-v0 | 25 | PPO | 14 | 3 (2–4) |
| FindImposterShape5-VLA-v0 | 25 | PPO | 15 | 3 (2–4) |
| FindImposterShape9-VLA-v0 | 25 | PPO | 15 | 3 (2–4) |
| Env ID | Horizon (steps) | Source | Median length | Gap |
|---|---|---|---|---|
| FindImposterShapeAndColor3x2-VLA-v0 | 25 | PPO | 16 | 3 (2–4) |
| FindImposterShapeAndColor3x3-VLA-v0 | 25 | PPO | 15 | 3 (2–4) |
| FindImposterShapeAndColor5x3-VLA-v0 | 40 | PPO | 15 | 3 (2–4) |
| Env ID | Horizon (steps) | Source | Median length | Gap |
|---|---|---|---|---|
| GatherAndRecall1-VLA-v0 | 200 | MP | 115 | undetermined |
| GatherAndRecall3-VLA-v0 | 400 | MP | 318 | undetermined |
| GatherAndRecall5-VLA-v0 | 600 | MP | 514 | undetermined |
| GatherAndRecall7-VLA-v0 | 800 | MP | 721 | undetermined |
| GatherAndRecall9-VLA-v0 | 1000 | MP | 900 | undetermined |
| Env ID | Horizon (steps) | Source | Median length | Gap |
|---|---|---|---|---|
| InterceptSlow-VLA-v0 | 60 | PPO | 43 | cue never hidden |
| InterceptMedium-VLA-v0 | 60 | PPO | 41 | cue never hidden |
| InterceptFast-VLA-v0 | 60 | PPO | 30 | cue never hidden |
| Env ID | Horizon (steps) | Source | Median length | Gap |
|---|---|---|---|---|
| InterceptGrabSlow-VLA-v0 | 60 | PPO | 24 | cue never hidden |
| InterceptGrabMedium-VLA-v0 | 60 | PPO | 33 | cue never hidden |
| InterceptGrabFast-VLA-v0 | 60 | PPO | 50 | cue never hidden |
| Env ID | Horizon (steps) | Source | Median length | Gap |
|---|---|---|---|---|
| RememberColor3-VLA-v0 | 25 | PPO | 15 | 3 (2–4) |
| RememberColor5-VLA-v0 | 25 | PPO | 15 | 3 (2–4) |
| RememberColor9-VLA-v0 | 25 | PPO | 15 | 3 (2–4) |
| RememberColor3-Long-VLA-v0 | 600 | MP | 337 | 250 (150–350) |
| RememberColor5-Long-VLA-v0 | 600 | MP | 390 | 250 (150–350) |
| RememberColor9-Long-VLA-v0 | 600 | MP | 378 | 250 (150–350) |
| Env ID | Horizon (steps) | Source | Median length | Gap |
|---|---|---|---|---|
| RememberShape3-VLA-v0 | 25 | PPO | 15 | 3 (2–4) |
| RememberShape5-VLA-v0 | 25 | PPO | 15 | 3 (2–4) |
| RememberShape9-VLA-v0 | 25 | PPO | 15 | 3 (2–4) |
| RememberShape3-Long-VLA-v0 | 600 | MP | 337 | 250 (150–350) |
| RememberShape5-Long-VLA-v0 | 600 | MP | 349 | 250 (150–350) |
| RememberShape9-Long-VLA-v0 | 600 | MP | 328 | 250 (150–350) |
| Env ID | Horizon (steps) | Source | Median length | Gap |
|---|---|---|---|---|
| RememberShapeAndColor3x2-VLA-v0 | 25 | PPO | 15 | 3 (2–4) |
| RememberShapeAndColor3x3-VLA-v0 | 25 | PPO | 14 | 3 (2–4) |
| RememberShapeAndColor5x3-VLA-v0 | 25 | PPO | 15 | 3 (2–4) |
| RememberShapeAndColor3x2-Long-VLA-v0 | 600 | MP | 312 | 250 (150–350) |
| RememberShapeAndColor3x3-Long-VLA-v0 | 600 | MP | 321 | 250 (150–350) |
| RememberShapeAndColor5x3-Long-VLA-v0 | 600 | MP | 322 | 250 (150–350) |
| Env ID | Horizon (steps) | Source | Median length | Gap |
|---|---|---|---|---|
| RotateLenientPos-VLA-v0 | 60 | PPO | 29 | cue never hidden |
| RotateLenientPosNeg-VLA-v0 | 60 | PPO | 19 | cue never hidden |
| Env ID | Horizon (steps) | Source | Median length | Gap |
|---|---|---|---|---|
| RotateStrictPos-VLA-v0 | 90 | PPO | 31 | cue never hidden |
| RotateStrictPosNeg-VLA-v0 | 90 | PPO | 22 | cue never hidden |
| Env ID | Horizon (steps) | Source | Median length | Gap |
|---|---|---|---|---|
| SeqOfColors3-VLA-v0 | 400 | MP | 165 | 3 (2–4) |
| SeqOfColors5-VLA-v0 | 400 | MP | 253 | 3 (2–4) |
| SeqOfColors7-VLA-v0 | 400 | MP | 342 | 3 (2–4) |
| SeqOfColors3-Long-VLA-v0 | 800 | MP | 559 | 225 (138–313) |
| SeqOfColors5-Long-VLA-v0 | 1000 | MP | 738 | 225 (138–313) |
| SeqOfColors7-Long-VLA-v0 | 1200 | MP | 916 | 225 (138–313) |
| Env ID | Horizon (steps) | Source | Median length | Gap |
|---|---|---|---|---|
| ShellGameColorLampTouch-VLA-v0 | 30 | PPO | 15 | 0 |
| Env ID | Horizon (steps) | Source | Median length | Gap |
|---|---|---|---|---|
| ShellGamePush-VLA-v0 | 30 | PPO | 15 | 0 |
| Env ID | Horizon (steps) | Source | Median length | Gap |
|---|---|---|---|---|
| ShellGameShuffleColorLampTouch-VLA-v0 | 60 | PPO | 44 | 28 (24–31) |
| ShellGameShuffleColorLampTouch-Long-VLA-v0 | 600 | MP | 331 | 250 (175–325) |
| Env ID | Horizon (steps) | Source | Median length | Gap |
|---|---|---|---|---|
| ShellGameShuffleTouch-VLA-v0 | 60 | PPO | 42 | 28 (24–31) |
| ShellGameShuffleTouch-Long-VLA-v0 | 600 | MP | 337 | 250 (175–325) |
| Env ID | Horizon (steps) | Source | Median length | Gap |
|---|---|---|---|---|
| ShellGameTouch-VLA-v0 | 30 | PPO | 22 | 0 |
| Env ID | Horizon (steps) | Source | Median length | Gap |
|---|---|---|---|---|
| TakeItBack-VLA-v0 | 60 | PPO | 27 | undetermined |
| Env ID | Horizon (steps) | Source | Median length | Gap |
|---|---|---|---|---|
| TimedTransferEasy-VLA-v0 | 200 | MP | 106 | 100 |
| TimedTransferMedium-VLA-v0 | 250 | MP | 152 | 150 |
| TimedTransferHard-VLA-v0 | 300 | MP | 203 | 200 |
| TimedTransferEasy-Long-VLA-v0 | 600 | MP | 298 | 300 |
| TimedTransferMedium-Long-VLA-v0 | 900 | MP | 488 | 500 |
| TimedTransferHard-Long-VLA-v0 | 1200 | MP | 962 | 1000 |
| Env ID | Horizon (steps) | Source | Median length | Gap |
|---|---|---|---|---|
| TraceShapeEasy-VLA-v0 | 250 | MP | 209 | 0 |
| TraceShapeMedium-VLA-v0 | 300 | MP | 210 | 0 |
| TraceShapeHard-VLA-v0 | 350 | MP | 207 | 0 |
| Env ID | Horizon (steps) | Source | Median length | Gap |
|---|---|---|---|---|
| TraceShapeSeqEasy-VLA-v0 | 1500 | MP | 702 | 0 |
| TraceShapeSeqMedium-VLA-v0 | 1500 | MP | 720 | 0 |
| TraceShapeSeqHard-VLA-v0 | 1500 | MP | 706 | 0 |
| Env ID | Memory type | Split | Horizon | Source | Mean len. | Median len. | Oracle succ. | Gap |
|---|---|---|---|---|---|---|---|---|
| ShellGameTouch-VLA-v0 | Spatial | Short | 30 | PPO | 21.5 | 22 | 1.000 | 0 |
| ShellGamePush-VLA-v0 | Spatial | Short | 30 | PPO | 15.0 | 15 | 1.000 | 0 |
| InterceptSlow-VLA-v0 | Spatial | Short | 60 | PPO | 43.8 | 43 | 1.000 | cue never hidden |
| InterceptMedium-VLA-v0 | Spatial | Short | 60 | PPO | 41.8 | 41 | 1.000 | cue never hidden |
| InterceptFast-VLA-v0 | Spatial | Short | 60 | PPO | 31.6 | 30 | 1.000 | cue never hidden |
| InterceptGrabSlow-VLA-v0 | Spatial | Short | 60 | PPO | 27.6 | 24 | 1.000 | cue never hidden |
| Field | Value |
|---|---|
| Initialization | pi05_base cross-embodiment checkpoint, fully fine-tuned |
| Codebase | openpi-comet [ Bai et al., 2025 ] , a JAX/Flax fork of openpi |
| Training data | LeRobot format, 14 tasks 250 episodes |
| (3,500 episodes, 798,296 frames) | |
| Image input | Two cameras stacked, each, upscaled to |
| Action space | pd_ee_delta_pose , 7 active dims, padded to 32 |
| Scale | Controlled memory difficulty | Environment & learning | |||||||
|---|---|---|---|---|---|---|---|---|---|
| Benchmark | Backend | Tasks | Types | Demos | Ret. | Load | H-strat. | Dense rew. | GPU vec. |
| MemoryBench [ Fang et al., 2025 ] | RLBench / Coppelia | 3 | 2 | 300 | ✗ | ✗ | ✗ | ✗ | ✗ |
| LIBERO-Mem [ Chung et al., 2026 ] | LIBERO / MuJoCo | 10 | 4 | 1,000 | ✗ | ✗ | |||
| RMBench [ Chen et al., 2026 ] | RoboTwin / SAPIEN | 9 | – | 450 | ✗ | ✗ | ✗ | ✗ | |
| RoboMME [ Dai et al., 2026 ] | ManiSkill / SAPIEN | 16 | 4 | 1,600 | ✗ | ✗ | ✗ | ✗ | ✗ |
| RoboMemArena [ Lei et al., 2026 ] | LIBERO / MuJoCo | 26 | 4 | 2,600 | ✗ | ✗ | ✗ | ✗ | ✗ |
| Task | No memory | VLA |
|---|---|---|
| InterceptFast | 0.00 | 0.29 |
| InterceptGrabFast | 0.00 | 0.00 |
| InterceptGrabMedium | 0.00 | 0.00 |
| InterceptGrabSlow | 0.00 | 0.00 |
| InterceptMedium | 0.36 | 0.55 |
| InterceptSlow | 0.05 | 0.08 |
| Loader | Regime | ShellGamePush | InterceptMedium | TakeItBack | RememberColor5 | RSC3x3 | Mean |
|---|---|---|---|---|---|---|---|
| A (shuffled) | exec-8 | 0.93 | 0.30 | 0.68 | 0.14 | 0.07 | 0.424 |
| A (shuffled) | exec-1 | 0.28 | 0.42 | 0.77 | 0.28 | 0.14 | 0.378 |
| B (episodic) | exec-8 | 0.95 | 0.31 | 0.77 | 0.18 | 0.03 | 0.448 |
| B (episodic) | exec-1 | 0.33 | 0.51 | 0.81 | 0.18 | 0.15 | 0.396 |
| Task | (16 steps, step 140,000) | (64 steps, best step 42,500) |
|---|---|---|
| BunchOfColors3 | 0.00 | 0.00 |
| ChainOfColors3 | 0.00 | 0.00 |
| GatherAndRecall3 | 0.24 | 0.26 |
| RememberColor3-Medium | 0.32 | 0.52 |
| RememberColor5-Medium | 0.08 | 0.54 |
| RememberShapeAndColor3x2-Medium | 0.18 | 0.44 |