We focus on language-conditioned flow-based manipulation, where robot flows (robot velocity fields) serve as embodiment-agnostic, motion-centric representations for leveraging data collected from multiple robot platforms. This task is crucial because language-conditioned manipulation is essential for practical robotic systems, yet scaling robot foundation models remains limited by the labor-intensive collection of embodiment-specific data. Existing methods either coarsely approximate robot flows with sparse keypoint displacements, or cannot handle language-conditioned manipulation. To address this limitation, we propose NarrativeFlow, which models robot flows as continuous velocity fields using a flow-matching formulation conditioned on language. Accordingly, NarrativeFlow generates robot flows that are physically consistent with real-world manipulation. To validate NarrativeFlow, we have conducted experiments on standard datasets for language-conditioned manipulation. The experimental results show that NarrativeFlow outperforms representative baseline methods on standard evaluation metrics. Furthermore, through real-world experiments, we show that NarrativeFlow achieves higher success rates than baseline methods across multiple manipulation tasks. The project page is available at https://shota0520.github.io/NarrativeFlow-project-page/
Figures & tables
Figure 1 : Overview of our approach to language-conditioned flow-based manipulation. We formulate robot flows as continuous velocity fields via flow matching. For training, we use data collected from multiple robot platforms. At test time, given an instruction and an initial image, our model generates a robot flow. The robot then executes the object manipulation task based on the robot flow.
Figure 2 : Architecture of NarrativeFlow. The vision-language fusion encoder integrates the embeddings of ℓ and I0 into the two subtask tokens: one for conditioning the Flow as Flow module to generate the robot flow, and the other for computing the Narrative Delta loss. During training, the Narrative Delta loss is used as an auxiliary training objective, aligning the corresponding subtask token with the narrative representations of task-relevant scene changes from I0 to the goal image.
Fractal
Bridge V2
Method
ADE ↓
FDE ↓
LTDR ↑ [%]
ADE ↓
FDE ↓
LTDR ↑ [%]
FLIP [ 13 ]
66.17
87.52
35.69
50.73
68.43
47.72
Im2Flow2Act [ 49 ]
37.14
47.74
60.61
51.48
70.93
47.97
Ours
21.68
31.59
76.29
30.59
42.42
66.30
Table 2 : Quantitative comparison between the proposed method and baseline methods. The best scores for each metric are shown in bold .
Figure 3 : Qualitative comparison between NarrativeFlow and a baseline method (Im2Flow2Act [ 49 ] ). In each panel, the first row presents ℓ , I0 , and the ground-truth robot flow. The second and third rows show the predicted robot flows of the baseline method and the proposed method, respectively.
Figure 4 : A failure case of NarrativeFlow on Bridge V2 [ 44 ] . The first row shows ℓ , I0 , and subsequent frames sampled from the ground-truth video in temporal order. The other rows show the robot flows of the ground truth, a baseline (Im2Flow2Act [ 49 ] ), and the proposed method in temporal order, respectively.
Model
Subtask [-0.2ex]tokens
Narrative Delta [-0.2ex]loss
Fractal
Bridge V2
ADE ↓
FDE ↓
LTDR ↑ [%]
ADE ↓
FDE ↓
LTDR ↑ [%]
(i)
0
✓
23.20
32.91
74.84
33.41
47.40
63.42
(ii)
1
✓
24.72
33.96
72.89
33.20
45.30
63.50
(iii)
2
25.83
35.72
71.87
33.79
47.81
63.00
(iv) Ours
2
✓
21.68
31.59
76.29
30.59
42.42
66.30
Table 3 : Quantitative results of ablation studies on the subtask tokens and the Narrative Delta loss. ✓ indicates the use of the Narrative Delta loss. The best scores for each metric are shown in bold .
Figure 5 : The Human Support Robot (HSR) used in our real-world experiments.
Method [%]
Mobile drawer closing
Mobile bin picking
Mobile cup stacking
Avg.
FLIP [ 13 ]
50
25
10
28
Im2Flow2Act [ 49 ]
65
45
15
42
Ours
85
60
20
55
Oracle
90
80
20
63
Table 4 : Quantitative results of real-world experiments. For each method, we report success rates across 20 trials per task, and the average over the three tasks. For the oracle method, the policy was conditioned on the corresponding ground-truth robot flows. The best scores among the proposed and baseline methods are shown in bold .
Figure 6 : Qualitative results of real-world experiments. The panels show successful executions of three manipulation tasks: (i) mobile drawer closing , (ii) mobile bin picking , and (iii) mobile cup stacking . In each panel, the top row shows ℓ , I0 , and the generated robot flow, while the bottom row shows the downstream manipulation conditioned on the robot flow.
Appendix figures & tables5 assets
Supplementary material from the paper’s appendix.
Appendix
Optimizer
AdamW
Batch size
128
Learning rate
1.0×10−4
LR schedule
Cosine decay
Warmup steps
9,000
Training steps
300,000
Weight decay
0.01
Appendix
Table A : Training settings of the proposed method.
Model
Nc
Loss [-0.2ex]function
Caption [-0.2ex]source
Fractal
Bridge V2
ADE ↓
FDE ↓
LTDR ↑ [%]
ADE ↓
FDE ↓
LTDR ↑ [%]
(i)
1
Cosine similarity
I0/G
21.70
31.68
76.26
31.24
43.64
65.42
(ii)
3
Cosine similarity
I0/G
21.71
31.68
76.24
30.63
42.80
66.09
(iii)
5
L2
I0/G
27.00
36.26
70.36
34.98
48.28
61.69
(iv)
5
Cosine similarity
Video
22.26
31.76
75.63
32.39
45.60
64.70
(v) Ours
5
Cosine similarity
I0/G
21.68
31.59
76.29
30.59
42.42
66.30
Appendix
Table B : Quantitative results of preliminary analysis of the Narrative Delta loss. I0 / G indicates that captions were generated from I0 and G , while Video indicates that captions were generated from multiple frames sampled from the video clip. The best scores for each metric are shown in bold .
Model
Caption generator
σ0
Fractal
Bridge V2
ADE ↓
FDE ↓
LTDR ↑ [%]
ADE ↓
FDE ↓
LTDR ↑ [%]
(i)
Qwen3.5-0.8B
0.05
26.40
36.59
71.07
32.72
44.60
63.89
(ii)
Qwen3.5-9B
0
26.43
36.83
70.86
34.35
48.01
62.28
(iii) Ours
Qwen3.5-9B
0.05
21.68
31.59
76.29
30.59
42.42
66.30
Appendix
Table C : Quantitative results of the additional ablation studies.
Method [%]
Mobile bin pushing
Mobile chair pushing
Mobile drawer closing
Mobile box closing
Mobile towel taking
Mobile bin picking
Mobile object placing
Mobile table bussing
Mobile laptop closing
Mobile water pouring
Mobile drawer opening
Mobile cup stacking
Mobile block stacking
Avg.
FLIP [ 13 ]
45
55
50
60
50
25
15
15
25
5
10
10
5
28
Im2Flow2Act [ 49 ]
70
70
65
65
60
45
50
40
45
20
20
15
20
45
Ours
90
90
85
85
70
60
65
70
50
35
20
20
20
58
Appendix
Table D : Additional quantitative results on real-world experiments across 13 manipulation tasks.