We focus on language-conditioned flow-based manipulation, where robot flows (robot velocity fields) serve as embodiment-agnostic, motion-centric representations for leveraging data collected from multiple robot platforms. This task is crucial because language-conditioned manipulation is essential for practical robotic systems, yet scaling robot foundation models remains limited by the labor-intensive collection of embodiment-specific data. Existing methods either coarsely approximate robot flows with sparse keypoint displacements, or cannot handle language-conditioned manipulation. To address this limitation, we propose NarrativeFlow, which models robot flows as continuous velocity fields using a flow-matching formulation conditioned on language. Accordingly, NarrativeFlow generates robot flows that are physically consistent with real-world manipulation. To validate NarrativeFlow, we have conducted experiments on standard datasets for language-conditioned manipulation. The experimental results show that NarrativeFlow outperforms representative baseline methods on standard evaluation metrics. Furthermore, through real-world experiments, we show that NarrativeFlow achieves higher success rates than baseline methods across multiple manipulation tasks. The project page is available at https://shota0520.github.io/NarrativeFlow-project-page/
Figures & tables
Figure 1 : Overview of our approach to language-conditioned flow-based manipulation. We formulate robot flows as continuous velocity fields via flow matching. For training, we use data collected from multiple robot platforms. At test time, given an instruction and an initial image, our model generates a robot flow. The robot then executes the object manipulation task based on the robot flow.
Figure 2 : Architecture of NarrativeFlow. The vision-language fusion encoder integrates the embeddings of ℓ and I0 into the two subtask tokens: one for conditioning the Flow as Flow module to generate the robot flow, and the other for computing the Narrative Delta loss. During training, the Narrative Delta loss is used as an auxiliary training objective, aligning the corresponding subtask token with the narrative representations of task-relevant scene changes from I0 to the goal image.
Fractal
Bridge V2
Method
ADE ↓
FDE ↓
LTDR ↑ [%]
ADE ↓
FDE ↓
LTDR ↑ [%]
FLIP [ 13 ]
66.17
87.52
35.69
50.73
68.43
47.72
Im2Flow2Act [ 49 ]
37.14
47.74
60.61
51.48
70.93
47.97
Ours
21.68
31.59
76.29
30.59
42.42
66.30
Table 2 : Quantitative comparison between the proposed method and baseline methods. The best scores for each metric are shown in bold .
Figure 3 : Qualitative comparison between NarrativeFlow and a baseline method (Im2Flow2Act [ 49 ] ). In each panel, the first row presents ℓ , I0 , and the ground-truth robot flow. The second and third rows show the predicted robot flows of the baseline method and the proposed method, respectively.
Figure 4 : A failure case of NarrativeFlow on Bridge V2 [ 44 ] . The first row shows ℓ , I0 , and subsequent frames sampled from the ground-truth video in temporal order. The other rows show the robot flows of the ground truth, a baseline (Im2Flow2Act [ 49 ] ), and the proposed method in temporal order, respectively.
Model
Subtask [-0.2ex]tokens
Narrative Delta [-0.2ex]loss
Fractal
Bridge V2
ADE ↓
FDE ↓
LTDR ↑ [%]
ADE ↓
FDE ↓
LTDR ↑ [%]
(i)
0
✓
23.20
32.91
74.84
33.41
47.40
63.42
(ii)
1
✓
24.72
33.96
72.89
33.20
45.30
63.50
(iii)
2
25.83
35.72
71.87
33.79
47.81
63.00
(iv) Ours
2
✓
21.68
31.59
76.29
30.59
42.42
66.30
Table 3 : Quantitative results of ablation studies on the subtask tokens and the Narrative Delta loss. ✓ indicates the use of the Narrative Delta loss. The best scores for each metric are shown in bold .
Figure 5 : The Human Support Robot (HSR) used in our real-world experiments.
Method [%]
Mobile drawer closing
Mobile bin picking
Mobile cup stacking
Avg.
FLIP [ 13 ]
50
25
10
28
Im2Flow2Act [ 49 ]
65
45
15
42
Ours
85
60
20
55
Oracle
90
80
20
63
Table 4 : Quantitative results of real-world experiments. For each method, we report success rates across 20 trials per task, and the average over the three tasks. For the oracle method, the policy was conditioned on the corresponding ground-truth robot flows. The best scores among the proposed and baseline methods are shown in bold .
Figure 6 : Qualitative results of real-world experiments. The panels show successful executions of three manipulation tasks: (i) mobile drawer closing , (ii) mobile bin picking , and (iii) mobile cup stacking . In each panel, the top row shows ℓ , I0 , and the generated robot flow, while the bottom row shows the downstream manipulation conditioned on the robot flow.
Appendix figures & tables5 assets
Supplementary material from the paper’s appendix.
Appendix
Optimizer
AdamW
Batch size
128
Learning rate
1.0×10−4
LR schedule
Cosine decay
Warmup steps
9,000
Training steps
300,000
Weight decay
0.01
Appendix
Table A : Training settings of the proposed method.
Model
Nc
Loss [-0.2ex]function
Caption [-0.2ex]source
Fractal
Bridge V2
ADE ↓
FDE ↓
LTDR ↑ [%]
ADE ↓
FDE ↓
LTDR ↑ [%]
(i)
1
Cosine similarity
I0/G
21.70
31.68
76.26
31.24
43.64
65.42
(ii)
3
Cosine similarity
I0/G
21.71
31.68
76.24
30.63
42.80
66.09
(iii)
5
L2
I0/G
27.00
36.26
70.36
34.98
48.28
61.69
(iv)
5
Cosine similarity
Video
22.26
31.76
75.63
32.39
45.60
64.70
(v) Ours
5
Cosine similarity
I0/G
21.68
31.59
76.29
30.59
42.42
66.30
Appendix
Table B : Quantitative results of preliminary analysis of the Narrative Delta loss. I0 / G indicates that captions were generated from I0 and G , while Video indicates that captions were generated from multiple frames sampled from the video clip. The best scores for each metric are shown in bold .
Model
Caption generator
σ0
Fractal
Bridge V2
ADE ↓
FDE ↓
LTDR ↑ [%]
ADE ↓
FDE ↓
LTDR ↑ [%]
(i)
Qwen3.5-0.8B
0.05
26.40
36.59
71.07
32.72
44.60
63.89
(ii)
Qwen3.5-9B
0
26.43
36.83
70.86
34.35
48.01
62.28
(iii) Ours
Qwen3.5-9B
0.05
21.68
31.59
76.29
30.59
42.42
66.30
Appendix
Table C : Quantitative results of the additional ablation studies.
Method [%]
Mobile bin pushing
Mobile chair pushing
Mobile drawer closing
Mobile box closing
Mobile towel taking
Mobile bin picking
Mobile object placing
Mobile table bussing
Mobile laptop closing
Mobile water pouring
Mobile drawer opening
Mobile cup stacking
Mobile block stacking
Avg.
FLIP [ 13 ]
45
55
50
60
50
25
15
15
25
5
10
10
5
28
Im2Flow2Act [ 49 ]
70
70
65
65
60
45
50
40
45
20
20
15
20
45
Ours
90
90
85
85
70
60
65
70
50
35
20
20
20
58
Appendix
Table D : Additional quantitative results on real-world experiments across 13 manipulation tasks.
Cross-embodiment data have become central to training robotic foundation models. To leverage such heterogeneous data, we focus on flow-based object manipulation, where robot flows (robot velocity fields) serve as embodiment-agnostic motion representations. Previous studies do not formulate robot flows as dense velocity fields, but as displacements of sparse keypoints, even though dense velocity fields better match the continuous-time nature of motions. To address this, we propose Flow as Flow, a framework that models robot flows as probability flows based on a flow matching formulation. By naturally modeling such velocity fields within this formulation, our method achieves efficient and high-quality robot flow generation. Across standard benchmarks, our method outperforms representative baseline methods on standard metrics, while achieving approximately 24× faster generation than standard flow matching. Furthermore, through real-world experiments evaluating 9 methods with 260 trials per method across 13 manipulation tasks, we show that our method achieves a higher average success rate than the baseline methods.
Generative vision-language-action policies have advanced robot manipulation, but they often exhibit instability under noise, partial observability, and stochastic initial conditions. During extended rollouts, small velocity errors accumulate, degrading execution reliability. Existing diffusion and flow-based policies typically assume homoscedastic residuals and lack explicit uncertainty modeling within action generation, limiting robustness during iterative rollout. We propose SUREFlow, a state-space uncertainty-aware residual flow matching framework built on a Mamba backbone. The method jointly predicts action velocities and input-dependent residual uncertainty, enabling selective refinement of unreliable action dimensions without environment feedback while preserving computational efficiency. On LIBERO, SUREFlow achieves 92.5% average success rate (SR), outperforming the Mamba-based MaIL by 34.2%. On LIBERO-PRO, it attains around 49% SR using only 179M parameters, achieving performance comparable to large VLAs with 3-7B parameters. SUREFlow source code is available on: https://github.com/tanvirnwu/SUREFlow
Md Tanvir Islam, Sai Navaneet Peddapalli, Sangmoon Lee +1
School of Electronic and Electrical Engineering, Kyungpook National University, Daegu 41566, Republic of Korea
World models offer a promising route toward robot planning by enabling agents to imagine and verify the consequences of actions before execution. However, current video-based world models often struggle to capture the physical constraints that govern manipulation, particularly contact. Further, their action conditioning is often constrained to specific embodiments such as parallel grippers. We propose \emph{Contact Flow}, an embodiment-agnostic action representation that encodes manipulation through the trajectory of 3D contact points between an actor and a target object. By discarding actor-specific appearance and kinematics, Contact Flow provides a shared conditioning signal for both human demonstrations and robotic execution. Therefore, we can train a large-scale video generative model on both human and robotic object interaction videos conditioned on Contact Flow, yielding a world model that predicts physically plausible manipulation outcomes. We integrate this model into a propose-imagine-verify-act pipeline, where generated rollouts are assessed by a vision-language model before execution. Experiments on the DROID dataset and real-world tabletop manipulation tasks demonstrate that Contact Flow enables transfer between human demonstrations and different robotic embodiments.