Organizations: School of Computer Science and Technology, Chongqing University of Posts and Telecommunications, Chongqing, China · Academy of Advanced Interdisciplinary Studies, Chongqing University of Posts and Telecommunications, Chongqing, China · School of Artificial Intelligence, Chongqing University of Posts and Telecommunications, Chongqing, China · ZTE Corporation, China · Towngas, China
Latent reasoning enables vision-language-action (VLA) models to transform multimodal observations into task-relevant internal states before generating continuous robot actions. While existing methods learn to generate or refine such states for each policy query, they discard successful reasoning after execution and therefore reconstruct similar computation from scratch. We present Reasoning and Flow Memory (FLOWMEM), a unified VLA model that turns successful latent computation into reusable reasoning experience. Rather than appending a fixed retrieved context, FLOWMEM dynamically retrieves and recomposes compatible latent fragments as the embodied context evolves, forming a reasoning route that follows the temporal structure and progress of successful computation. The route is then refined using current visual and proprioceptive evidence before it conditions action generation. Experiments on RoboMME and LIBERO-Plus show that FLOWMEM attains 48.0% and 77.3% success, outperforming memory-free policies by 1.7 and 4.1 percentage points, respectively. These results demonstrate the value of reusing successful latent computation for closed-loop VLA control.
Figures & tables
Figure 1: Comparison between transient latent reasoning and the proposed reusable reasoning-flow formulation.
Memory
Method
Counting
Permanence
Reference
Imitation
Avg.
None
π0.5
28.8
17.0
17.2
8.8
17.9
Action history
π0.5 + past actions
29.1
22.8
15.9
11.2
19.7
Symbolic
SimpleSG + QwenVL
44.6
19.6
25.2
26.6
29.0
GroundSG + QwenVL
38.0
39.3
31.6
21.9
32.7
Episodic
MemER
48.8
53.2
38.0
29.5
42.4
Perceptual
SAM2Act+
35.3
26.0
16.8
7.3
21.4
Table 1: Main results on RoboMME. Success rate (%) under the official Full-16 protocol. Higher is better. Best and second-best are bolded and underlined.
Method
Venue
Spatial
Object
Goal
Long
Avg.
OpenVLA
CoRL’24
19.4
14.0
15.1
14.3
15.6
UniVLA
RSS’25
55.5
36.7
40.7
39.9
42.9
OpenVLA-OFT
RSS’25
84.0
66.5
63.0
66.4
69.6
MemoryVLA
ICLR’26
64.7
52.4
50.7
52.8
55.0
MergeVLA
CVPR’26
83.7
80.4
65.6
59.0
72.0
VLA-Adapter
AAAI’26
44.7
40.8
43.6
47.9
44.2
Table 2: Main results on LIBERO-Plus. Success rate (%) following suite-specific standard-LIBERO training and zero-shot LIBERO-Plus evaluation. Avg. pools all 10,030 instances across the four suites. Higher is better. Best and second-best are bolded and underlined.
Method
Flow
Multi-src.
Refine
RoboMME
LIBERO-Plus
LaST 0
46.3
73.2
Single-source FlowMem
✓
✓
44.0
76.5
FlowMem w/o refinement
✓
✓
46.3
76.7
FlowMem (Ours)
✓
✓
✓
48.0
77.3
Table 5: Core ablation of FlowMem. Variants share the same action interface and latent-token budget. Higher is better. Best and second-best are bolded and underlined.
Appendix figures & tables3 assets
Supplementary material from the paper’s appendix.
Appendix
Method
BinFill
PickX
SwingX
StopCube
V-Unmask
V-Unmask-S
B-Unmask
B-Unmask-S
FrameSamp+Modul
39.6
87.3
92.0
42.0
32.7
24.4
25.1
18.2
LaST 0
46.0
96.0
90.0
54.0
34.0
24.0
26.0
28.0
FlowMem (Ours)
50.0
98.0
96.0
56.0
40.0
26.0
24.0
26.0
Appendix
Table 6: Per-task RoboMME results. Success rate (%) under the official 16-task, 50-episode-per-task protocol. Higher is better. Best results are bolded.
Method
Venue
Camera
Robot
Language
Light
Background
Noise
Layout
Total
OpenVLA
CoRL’24
0.8
3.5
23.0
8.1
34.8
15.2
28.5
15.6
UniVLA
RSS’25
1.8
46.2
69.6
69.0
81.0
21.2
31.9
42.9
OpenVLA-OFT
RSS’25
56.4
31.9
79.5
88.7
93.3
75.8
74.2
69.6
MemoryVLA
ICLR’26
10.3
50.2
79.5
65.3
89.4
31.1
75.1
55.0
MergeVLA
CVPR’26
61.7
44.8
75.7
92.0
93.0
73.7
75.1
72.0
VLA-Adapter
AAAI’26
6.1
29.1
66.2
56.5
70.5
25.7
69.2
44.2
Appendix
Table 7: Perturbation-wise LIBERO-Plus results. Success rate (%) following suite-specific standard-LIBERO training and zero-shot LIBERO-Plus evaluation. Each perturbation column pools its instances across the four suites; Total pools all 10,030 instances. Higher is better. Best and second-best are bolded and underlined.
Method
Time / episode (s)
SR (%)
LaST 0
19.6
46.3
FlowMem
15.5
48.0
Appendix
Table 8: Inference efficiency on RoboMME. Time / episode is the mean total online control time from observation to action-chunk return over the profiling episodes. Lower time and higher SR are better.
Mainstream Vision-Language-Action (VLA) models predict actions primarily from the current observation under a Markovian assumption, thus struggling with long-horizon, temporally dependent tasks. Existing memory-augmented VLAs either expand the observation window or retrieve history from the memory bank as auxiliary policy-side context. However, they leave memory outside the native latent embedding space of VLA reasoning, preventing historical experience from being fluidly interleaved with multimodal reasoning and action formation. To this end, we introduce LaMem-VLA, a latent-memory-native framework that reconstructs historical experience into latent memory tokens and directly interweaves them with VLA reasoning. At its core, LaMem-VLA introduces four coordinated components: (i) a curator that organizes historical experience into two complementary short-term and long-term memory vaults; (ii) a seeker that queries both vaults using the multimodal cognition to retrieve context-relevant evidence; (iii) a condenser that reconstructs the retrieved evidence into compact short-term and long-term latent memory tokens; and (iv) a weaver that injects these memory tokens with the current observation and instruction into one continuous embedding sequence. By representing, retrieving, and consuming historical experience entirely in the same continuous latent space, LaMem-VLA enables memory to directly participate in VLA reasoning and guide action generation under a bounded context. Extensive experiments on SimplerEnv and LIBERO demonstrate the superiority of our LaMem-VLA.
Hongyu Qu, Jianzhe Gao, Xiaobin Hu +6
Nanjing University of Science and Technology · Zhejiang University · National University of Singapore
Natural language is a powerful reasoning medium for language and vision-language models, but it is mismatched to the granularity of continuous control. Text and explicit subgoals operate at task-level granularity, whereas vision-language-action (VLA) policies must choose actions at a much finer temporal scale; a single reasoning step can therefore span many action chunks while remaining only weakly coupled to the action needed now. This suggests a different question for VLA: what should play the role of language? We argue that a useful VLA reasoning medium must be shareable across model instances, verifiable through downstream action improvement, and aligned with temporally extended control structure. Based on this view, we propose Continuous Reasoning for Vision-Language-Action. Our model first predicts continuous reasoning in the form of a structured set of continuous thoughts, then reuses them as shared context for chunk-structured action generation. Better action prediction alone does not certify good reasoning: if the same internal medium cannot be shared across model instances and independently verified through improved downstream control, the added latent may simply become a model-private shortcut that helps on seen behaviors without supporting generalizable control. We therefore instantiate continuous reasoning as a shared Gaussian latent interface and train it with a self-verification objective in which an exponential-moving-average teacher must successfully consume the student's reasoning when predicting target actions. Empirically, Continuous Reasoning improves LIBERO-PRO robustness and performs strongly on real robots, raising mean subtask success over π0.5 by 40.4% on TX-G2, an AgiBot G2-compatible variant, and 26.3% on HSR. This suggests that reasoning in VLA is less about extra tokens than about a shareable, verifiable internal language for action.
Vision-language-action (VLA) models remain constrained by the scarcity of action-labeled robot data, whereas action-free videos provide abundant evidence of how the physical world changes. Latent action models offer a promising way to extract such priors from videos, but reconstruction-trained latent codes are not necessarily suitable for policy generation: they may predict future observations while lacking the structure needed to be reused or generated coherently with robot actions. We introduce ALAM (Algebraic Latent Action Model), an Algebraically Consistent Latent Action Model that turns temporal relations in action-free video into structural supervision. Given frame triplets, ALAM learns latent transitions that are grounded by reconstruction while being regularized by composition and reversal consistency, encouraging a locally additive transition space. For downstream VLA learning, we freeze the pretrained encoder and use its latent transition sequences as auxiliary generative targets, co-generated with robot actions under a joint flow-matching objective. This couples structured latent transitions with flow-based policy generation, allowing the policy to exploit ALAM's locally consistent transition geometry without requiring latent-to-action decoding. Representation probes show that ALAM reduces additivity and reversibility errors by 25-85 times over unstructured latent-action baselines and improves long-horizon cumulative reconstruction. When transferred to VLA policies, ALAM raises the average success rate from 47.9% to 85.0% on MetaWorld MT50 and from 94.1% to 98.1% on LIBERO, with consistent gains on real-world manipulation tasks. Ablations further confirm that the strongest improvements arise from the synergy between algebraically structured latent transitions and joint flow matching.
Zuojin Tang, Haoyun Liu, Xinyuan Chang +11
Zhejiang University · Amap, Alibaba Group · Nanjing University +5