Organizations: School of Computer Science and Technology, Chongqing University of Posts and Telecommunications, Chongqing, China · Academy of Advanced Interdisciplinary Studies, Chongqing University of Posts and Telecommunications, Chongqing, China · School of Artificial Intelligence, Chongqing University of Posts and Telecommunications, Chongqing, China · ZTE Corporation, China · Towngas, China
Latent reasoning enables vision-language-action (VLA) models to transform multimodal observations into task-relevant internal states before generating continuous robot actions. While existing methods learn to generate or refine such states for each policy query, they discard successful reasoning after execution and therefore reconstruct similar computation from scratch. We present Reasoning and Flow Memory (FLOWMEM), a unified VLA model that turns successful latent computation into reusable reasoning experience. Rather than appending a fixed retrieved context, FLOWMEM dynamically retrieves and recomposes compatible latent fragments as the embodied context evolves, forming a reasoning route that follows the temporal structure and progress of successful computation. The route is then refined using current visual and proprioceptive evidence before it conditions action generation. Experiments on RoboMME and LIBERO-Plus show that FLOWMEM attains 48.0% and 77.3% success, outperforming memory-free policies by 1.7 and 4.1 percentage points, respectively. These results demonstrate the value of reusing successful latent computation for closed-loop VLA control.
Figures & tables
Figure 1: Comparison between transient latent reasoning and the proposed reusable reasoning-flow formulation.
Memory
Method
Counting
Permanence
Reference
Imitation
Avg.
None
π0.5
28.8
17.0
17.2
8.8
17.9
Action history
π0.5 + past actions
29.1
22.8
15.9
11.2
19.7
Symbolic
SimpleSG + QwenVL
44.6
19.6
25.2
26.6
29.0
GroundSG + QwenVL
38.0
39.3
31.6
21.9
32.7
Episodic
MemER
48.8
53.2
38.0
29.5
42.4
Perceptual
SAM2Act+
35.3
26.0
16.8
7.3
21.4
Table 1: Main results on RoboMME. Success rate (%) under the official Full-16 protocol. Higher is better. Best and second-best are bolded and underlined.
Method
Venue
Spatial
Object
Goal
Long
Avg.
OpenVLA
CoRL’24
19.4
14.0
15.1
14.3
15.6
UniVLA
RSS’25
55.5
36.7
40.7
39.9
42.9
OpenVLA-OFT
RSS’25
84.0
66.5
63.0
66.4
69.6
MemoryVLA
ICLR’26
64.7
52.4
50.7
52.8
55.0
MergeVLA
CVPR’26
83.7
80.4
65.6
59.0
72.0
VLA-Adapter
AAAI’26
44.7
40.8
43.6
47.9
44.2
Table 2: Main results on LIBERO-Plus. Success rate (%) following suite-specific standard-LIBERO training and zero-shot LIBERO-Plus evaluation. Avg. pools all 10,030 instances across the four suites. Higher is better. Best and second-best are bolded and underlined.
Method
Flow
Multi-src.
Refine
RoboMME
LIBERO-Plus
LaST 0
46.3
73.2
Single-source FlowMem
✓
✓
44.0
76.5
FlowMem w/o refinement
✓
✓
46.3
76.7
FlowMem (Ours)
✓
✓
✓
48.0
77.3
Table 5: Core ablation of FlowMem. Variants share the same action interface and latent-token budget. Higher is better. Best and second-best are bolded and underlined.
Appendix figures & tables3 assets
Supplementary material from the paper’s appendix.
Appendix
Method
BinFill
PickX
SwingX
StopCube
V-Unmask
V-Unmask-S
B-Unmask
B-Unmask-S
FrameSamp+Modul
39.6
87.3
92.0
42.0
32.7
24.4
25.1
18.2
LaST 0
46.0
96.0
90.0
54.0
34.0
24.0
26.0
28.0
FlowMem (Ours)
50.0
98.0
96.0
56.0
40.0
26.0
24.0
26.0
Appendix
Table 6: Per-task RoboMME results. Success rate (%) under the official 16-task, 50-episode-per-task protocol. Higher is better. Best results are bolded.
Method
Venue
Camera
Robot
Language
Light
Background
Noise
Layout
Total
OpenVLA
CoRL’24
0.8
3.5
23.0
8.1
34.8
15.2
28.5
15.6
UniVLA
RSS’25
1.8
46.2
69.6
69.0
81.0
21.2
31.9
42.9
OpenVLA-OFT
RSS’25
56.4
31.9
79.5
88.7
93.3
75.8
74.2
69.6
MemoryVLA
ICLR’26
10.3
50.2
79.5
65.3
89.4
31.1
75.1
55.0
MergeVLA
CVPR’26
61.7
44.8
75.7
92.0
93.0
73.7
75.1
72.0
VLA-Adapter
AAAI’26
6.1
29.1
66.2
56.5
70.5
25.7
69.2
44.2
Appendix
Table 7: Perturbation-wise LIBERO-Plus results. Success rate (%) following suite-specific standard-LIBERO training and zero-shot LIBERO-Plus evaluation. Each perturbation column pools its instances across the four suites; Total pools all 10,030 instances. Higher is better. Best and second-best are bolded and underlined.
Method
Time / episode (s)
SR (%)
LaST 0
19.6
46.3
FlowMem
15.5
48.0
Appendix
Table 8: Inference efficiency on RoboMME. Time / episode is the mean total online control time from observation to action-chunk return over the profiling episodes. Lower time and higher SR are better.