Learning from demonstration has enabled impressive robot behaviors. A common choice for policy learning is to use diffusion or flow matching (Flow-Policies), which often outperforms direct action regression trained with mean squared error (MSE-Policies). This gap is commonly attributed to multimodal demonstrations. We revisit this gap from the perspective of statistical modeling: how action-prediction residuals shape policy optimization. Our analysis of real-world robot demonstration data reveals substantial state-dependent variation in residual scales and heavier-than-Gaussian tails. While both MSE-Policies and Flow-Policies exhibit heavy-tailed action residuals, their training gradients behave differently: MSE allocates more gradient magnitude to observations with large action residuals, which hurts optimization. Motivated by these findings, we introduce heteroscedastic Student-t action regression (HT-Policies), which learns input-dependent residual scales and reduces the influence of heavy tails. HT-Policies predict action chunks with a single feed-forward pass and can reuse pretrained flow-matching-based policy networks as the backbone. Across four simulation benchmarks and real-robot evaluations, HT-Policies achieves success rates competitive with generative policy baselines, both when trained from scratch and from pretrained vision-language-action and world-action models, despite being faster in training and inference. Together, these findings shed light on the practical advantages of generative objectives in robot learning from demonstrations and offer an efficient direct-regression alternative for a range of architectures and tasks. Project page: https://the-labone.github.io/regression-policy-project/
Figures & tables
Figure 1: (a) We attribute the performance gap between MSE- and Flow-Policies to different statistical modeling of prediction residuals. (b) We found observation-dependent and heavy-tailed residuals for MSE-Policies on multiple benchmarks. (c) Benchmark success rates (%) for different policy classes. We use flow-matching for RoboCasa-GR1 and SIMPLER, and diffusion for RoboMimic.
Figure 2: Example tasks. We show 4 tasks from RoboCasa-GR1, Bridge, Fractal, and Tool-Hang.
Figure 3: Heavy-tailed residuals. Both MSE-Policies (a) and Flow-Policies (b) exhibit heavy-tailed action-prediction residuals. Percentages report scalar residuals exceeding three times their coordinate RMS in absolute value.
Figure 4: Gradient distribution for different objectives and noise levels. We show the relative gradient contribution (%) of each action-residual decile. For Flow, gradient RMS is computed separately for high-noise ( t<0.5 ) and low-noise ( t>0.5 ) regimes. See also the Appendix Figure 12 .
Figure 5: Input-dependent residual scale. We evaluate a GR00T N1.7 model trained with the MSE regression loss. (a) Residual vs. action RMS. (b) Residual RMS vs. episode progress on three Bridge tasks. Shaded bands indicate 95% confidence intervals..
Figure 6: Heavy tails after normalization. (a) HG-regression residuals normalized by predicted σ(o) ; (b) flow-matching residuals normalized by per-observation action sampling.
Backbone
DP
MSE
HT
U-Net
93.8/84.7
81.7/72.3
95.3 / 85.6
Transformer
94.3 / 84.1
82.6/74.6
91.9/82.9
Table 1: Success rates (%) for policies trained from scratch. (a) RoboMimic. Best / mean over the last ten evaluated checkpoints ( Chi et al., 2023 ) , averaged over nine tasks and three seeds per task, with state and image results weighted equally. (b) RoboCasa-GR1. A randomly initialized GR00T N1.7 action head trained with the official fine-tuning recipe.
Figure 7: Real-robot success rates (%). We evaluate on 3 tasks: Push-T, Insert-T, and Peel-Note.
Model
Benchmark
Flow (Official)
Our Evaluations (10 seeds)
Flow
MSE
HT
π0.5
LIBERO (4-suite avg.)
96.9
97.09±0.04
96.72±0.09
96.97±0.07
GR00T N1.7
RoboCasa-GR1
44.5
42.34±0.68
35.55±0.67
50.01±0.51
SIMPLER / Bridge
62.3
62.06±0.40
58.51±0.37
62.31±0.65
SIMPLER / Fractal
72.5
67.83±0.39
60.60±0.43
65.25±0.44
Cosmos 3
LIBERO-10
95.2
95.60±0.90
95.10±0.26
96.50±0.45
Table 2: Success rate (%) for fine-tuning from pretrained policies. We report the mean success rate and its standard error of a single checkpoint across 10 evaluation seeds. We include official results for the flow-matching–based fine-tuning for reference. For π0.5 , we report the average performance in four LIBERO suites. Bold marks highlight the best-performing model in each row.
Figure 8: Residual modeling and optimization. Top: Policy success rates across four settings. Bottom: Each action-residual decile’s share of summed per-sample parameter-gradient RMS under MSE, HT, and low-noise Flow ( t>0.5 ). Each decile contains approximately 10% of the samples.
Figure 9: Mean log-likelihood vs. policy success rate.
Figure 10: Training curves for different models on different benchmarks. HT-Policies reach strong task success rates with fewer training steps than Flow-Policies.
GR00T N1.7
π0.5
Cosmos3-Nano
Metric
Flow
HT
Speed-up
Flow
HT
Speed-up
Flow
HT
Speed-up
Head latency ↓
29.30
7.92
3.70×
62.50
6.42
9.74×
1266.10
65.96
19.19×
Total latency ↓
46.20
24.52
1.88×
123.58
67.94
1.82×
1271.76
71.34
17.83×
FPS ↑
21.65
40.78
1.88×
8.09
14.72
1.82×
0.79
14.02
17.83×
Table 3: Inference efficiency benchmarking. We evaluate all policies on a single NVIDIA RTX 5090 GPU with batch size 1. Latencies are in milliseconds.
Figure 11: Selection of ν . (a) Success rate vs. ν on RoboCasa-GR1 and (b) Success rate vs. c=ν/d on three datasets, where d is the action-chunk dimension. ν→∞ gives heteroscedastic Gaussian.
Appendix figures & tables30 assets
Supplementary material from the paper’s appendix.
Appendix
Dataset
Episodes
Observations
d
MSE/HG coordinates
RoboCasa-GR1
240
960
232
222,720
Bridge
400
1,600
48
76,800
Fractal
400
1,600
48
76,800
Tool-Hang
200
1,600
144
230,400
Appendix
Table 4: Residual-tail samples. Bridge, Fractal, and Tool-Hang exclude the gripper channel from the plotted residuals. Each Flow-Policy contributes 64 independently evaluated action draws per observation.
Figure 12: Gradient imbalance across objectives and noise levels. Curves show the relative gradient contribution (%) of each action-residual decile. For flow matching, gradient RMS is averaged over noise draws and timesteps separately for high noise ( t<0.5 ) and low noise ( t>0.5 ).
Setting
MSE
Flow
HT
RoboCasa-GR1
1,536
1,536
1,536
Bridge
2,048
2,048
2,048
Fractal
2,048
2,048
2,048
Tool-Hang
4,096
4,096
4,096
Appendix
Table 5: Observation counts for the parameter-gradient diagnostics. GR00T is trained for 60k updates on GR1 and 20k updates on Bridge and Fractal. Tool-Hang uses state observations and Chi-UNet.
Dataset
Quantile function Q(u)
High noise ( t<0.5 )
Low noise ( t>0.5 )
RoboCasa-GR1, Bridge, Fractal
0.999[1−(1−u)2/3]
Points 1–10 ( 62.5% ) [0,0.4795)
Points 11–16 ( 37.5% ) [0.4795,0.999]
Tool-Hang
u
Points 1–8 ( 50% ) [0,0.5)
Points 9–16 ( 50% ) [0.5,1]
Appendix
Table 6: Flow-matching timestep groups for gradient diagnostics. Ranges span the full quantile intervals represented by each group; percentages denote their training probability mass.
Figure 13: Gaussian tail-bound comparison across datasets. GR00T uses 60k updates on RoboCasa-GR1 and 20k on Bridge and Fractal; LIBERO results use π0.5 . Light curves show individual timesteps, the thick blue curve their pointwise maximum, and the dashed curve the Gaussian norm reference.
Figure 14: Gaussian tail-bound comparison across training stages. GR00T on RoboCasa-GR1 after 5k, 10k, 20k, 40k, and 60k updates, with d=232 . Curves follow the same convention as Figure 13 .
Figure 15: Gradient allocation and policy success across training stages on RoboCasa-GR1. Columns show results after 5k, 10k, 20k, 40k, and 60k updates. Top: Each action-residual decile’s share of summed per-sample parameter-gradient RMS, with flow-matching gradients evaluated at low-noise timesteps ( t>0.5 ). The dashed line marks the 10% sample share of each decile. Bottom: Success rates of MSE-Policies, Flow-Policies, and HT-Policies at the corresponding training stages; HT-Policies have the highest success rate at all five stages.
Model / benchmark
K=1
K=16 mean
Δ (pp)
GR00T N1.7 / RoboCasa-GR1
40.63
43.23
+2.61
π0.5 / LIBERO
97.250
97.325
+0.075
Cosmos3-Nano / LIBERO-10
96.00
98.00
+2.00
Appendix
Table 7: Action averaging in fixed Flow-Policies. Success rates (%) from one evaluation run per condition. Δ is the change in percentage points, computed before rounding.
Task
K=1
K=16 mean
Δ (pp)
PnPBottle → CabinetClose
8/20 (40.00)
12/20 (60.00)
+20.00
PnPCan → DrawerClose
2/20 (10.00)
6/20 (30.00)
+20.00
PnPCup → DrawerClose
0/20 (0.00)
4/20 (20.00)
+20.00
PnPMilk → MicrowaveClose
3/20 (15.00)
1/20 (5.00)
−10.00
PnPPotato → MicrowaveClose
4/20 (20.00)
1/20 (5.00)
−15.00
PnPWine → CabinetClose
3/20 (15.00)
3/21 (14.29)
−0.71
Appendix
Table 8: Action averaging on RoboCasa-GR1. Entries give successes / completed rollouts, with success rates (%) in parentheses. Task names follow Table 22 . The mean weights all 24 tasks equally; Δ is computed before rounding.
LIBERO-10
LIBERO-Goal
LIBERO-Object
LIBERO-Spatial
ID
K=1
K=16
K=1
K=16
K=1
K=16
K=1
K=16
0
96/100
93/100
99/100
99/100
100/100
100/100
100/100
100/100
1
98/100
98/100
98/100
100/100
100/100
100/100
100/100
99/100
2
99/100
96/100
89/100
91/100
100/100
100/100
100/100
100/100
3
100/100
100/100
94/100
94/100
97/100
98/100
97/100
100/100
4
99/100
98/100
100/100
100/100
100/100
99/100
98/100
98/100
Appendix
Table 9: Action averaging for π0.5 on all four LIBERO suites. Entries give successes / completed rollouts. Task IDs are listed in Tables 20 and 21 .
ID
Task
K=1
K=16 mean
0
Put both the alphabet soup and the tomato sauce in the basket.
8/10
9/10
1
Put both the cream cheese box and the butter in the basket.
10/10
10/10
2
Turn on the stove and put the moka pot on it.
9/10
10/10
3
Put the black bowl in the bottom drawer of the cabinet and close it.
10/10
10/10
4
Put the white mug on the left plate and the yellow and white mug on the right plate.
10/10
10/10
5
Pick up the book and place it in the back compartment of the caddy.
10/10
10/10
Appendix
Table 10: Action averaging for Cosmos3-Nano on LIBERO-10. Entries give successes / completed rollouts. Task IDs follow the LIBERO-10 ordering in Table 21 .
Method
α
pass@1
pass@2
pass@4
pass@8
Flow-Policy (native)
—
76.25
93.68
99.67
100.00
HT-Policy (no noise)
0
81.00
—
—
—
HT-Policy ( σ -scaled)
0.125
79.38
95.61
99.76
100.00
HT-Policy (fixed scale)
0.125
78.88
95.39
99.81
100.00
HT-Policy ( σ -scaled)
0.25
78.63
95.82
99.90
100.00
HT-Policy (fixed scale)
0.25
76.88
94.25
99.56
100.00
Appendix
Table 11: Exploration pass@ k (%) on state-based Tool-Hang with Chi-Transformer. Estimates are computed per initial state from eight rollouts and averaged over 100 states. α controls HT-Policy sampling noise. Dashes indicate unreported values or an inapplicable noise multiplier.
Model / dataset
Updates
Batch
Base LR
Warmup
ν
π0.5 / LIBERO
30,000
32
2.5×10−5
1,000
1,400
GR00T / RoboCasa-GR1
60,000
512
10−4
3,000
928
GR00T / Bridge
20,000
1,024
10−4
1,000
224
GR00T / Fractal
20,000
1,024
10−4
1,000
224
Cosmos3-Nano / LIBERO-10
2,000
≤2,048
5×10−5
500
640
Appendix
Table 12: Training budgets for the large-model experiments. Batch size is global. Cosmos3-Nano batch size depends on packing.
Dataset
Episodes
Frames
H
Da
d
LIBERO, four suites ( π0.5 )
1,693
273,465
50
7
350
RoboCasa-GR1
24,000
5,820,277
8
29
232
Bridge
53,192
1,893,026
8
7
56
Fractal
87,212
3,786,400
8
7
56
LIBERO-10 (Cosmos3-Nano)
379
101,469
16
10
160
Appendix
Table 13: Training data and effective action dimensions. Episode and frame counts refer to the datasets used for training.
Model / benchmark
Tasks
Rollout budget/ task/seed
Execute
Flow steps
Step limit
π0.5 / Spatial and Object
2×10
100
10
10
280
π0.5 / Goal
10
100
10
10
300
π0.5 / LIBERO-10
10
100
10
10
520
GR00T / RoboCasa-GR1
24
20
8
4
720
GR00T / Bridge–WidowX
7
50
4
4
300
GR00T / Fractal–Google Robot
6
100
1
4
300
Appendix
Table 14: Evaluation settings for the 10-seed comparison in Table 2 . Rollout budgets are per task and evaluation seed; actual GR00T rollout counts can slightly exceed these budgets. “Execute” is the number of actions applied before replanning. MSE-Policies and HT-Policies each use a single forward pass.
U-Net
Transformer
Task
DP-C
HT
DP-T
HT
best/last10
best/last10
best/last10
best/last10
Lift-PH
1.00 /0.98
1.00 / 1.00
1.00 / 1.00
1.00 /0.93
Lift-MH
1.00 /0.97
1.00 / 1.00
1.00 / 1.00
1.00 /0.96
Can-PH
1.00 /0.96
1.00 / 1.00
1.00 / 1.00
1.00 /0.94
Can-MH
1.00 /0.96
1.00 / 0.98
1.00 /0.94
1.00 / 0.98
Appendix
Table 15: DP and HT with state observations. Bold marks the highest value per row and metric across both backbones, before rounding.
U-Net
Transformer
Task
DP-C
HT
DP-T
HT
best/last10
best/last10
best/last10
best/last10
Lift-PH
1.00 / 1.00
1.00 /0.95
1.00 / 1.00
1.00 / 1.00
Lift-MH
1.00 / 1.00
1.00 /0.95
1.00 /0.99
1.00 /0.99
Can-PH
1.00 /0.97
1.00 /0.94
1.00 /0.98
1.00 / 1.00
Can-MH
1.00 /0.96
1.00 /0.92
1.00 / 0.98
1.00 /0.95
Appendix
Table 16: DP and HT with image observations. Bold marks the highest value per row and metric across both backbones, before rounding.
Method
Lift
Can
Square
Transport
Tool-Hang
Push-T
Mean
PH
MH
PH
MH
PH
MH
PH
MH
Sudeep-DiT
Flow
1.00 / 1.00
1.00 /0.99
1.00 / 1.00
1.00 /0.94
1.00 / 0.94
0.88/0.75
0.80/0.70
0.40/0.27
0.86/0.75
0.98 /0.95
0.89/0.83
MSE
1.00 / 1.00
1.00 /0.99
1.00 /0.98
0.92/0.90
0.94/0.86
0.72/0.53
0.50/0.44
0.12/0.06
0.52/0.39
0.92/0.83
0.76/0.70
Straight Flow
1.00 / 1.00
1.00 /0.98
1.00 /0.99
0.96/0.90
0.96/0.93
0.72/0.66
0.56/0.48
0.20/0.14
0.70/0.59
0.90/0.86
0.80/0.75
MIP
1.00 / 1.00
1.00 /0.99
1.00 / 1.00
0.98/0.95
0.98/ 0.94
0.90/0.81
0.76/0.68
0.44/0.38
0.92 / 0.88
0.95/0.92
0.89/0.86
Appendix
Table 17: Scores on state-based tasks. Entries report best/last5, averaged over three seeds. RoboMimic reports success; Push-T reports normalized coverage. Baselines follow Table 11 of Pan et al. (2026) . Bold marks the best value per task and metric across all methods and backbones, before rounding.
Method
Lift
Can
Square
Transport
Tool-Hang
Push-T
Mean
PH
MH
PH
MH
PH
MH
PH
MH
Sudeep-DiT
Flow
1.00 / 1.00
1.00 / 1.00
1.00 /0.99
0.96/0.94
0.96/ 0.94
0.82/0.76
0.84/0.83
0.32/0.20
0.78 /0.57
0.92 / 0.89
0.86/0.81
MSE
1.00 / 1.00
1.00 /0.99
1.00 / 1.00
0.92/0.81
0.94/0.84
0.74/0.67
0.74/0.56
0.14/0.08
0.28/0.18
0.83/0.77
0.76/0.69
Straight Flow
1.00 /0.99
1.00 /0.99
1.00 /0.98
0.98/0.95
1.00 /0.93
0.82/0.72
0.86/0.83
0.26/0.19
0.46/0.40
0.85/0.79
0.82/0.78
MIP
1.00 / 1.00
1.00 /0.99
1.00 /0.98
1.00 /0.96
1.00 /0.92
0.90/0.83
0.90/0.84
0.50/0.31
0.76/ 0.66
0.91/0.87
0.90/0.84
Appendix
Table 18: Scores on image-based tasks. Entries report best/last5, averaged over three seeds. RoboMimic reports success; Push-T reports normalized coverage. Baselines follow Table 12 of Pan et al. (2026) . Bold marks the best value per task and metric across all methods and backbones, before rounding.
Suite
Flow (OpenPI)
Flow
HT
LIBERO-Spatial
98.80
97.13±0.11
97.71±0.13
LIBERO-Object
98.20
99.44±0.06
99.80±0.00
LIBERO-Goal
98.00
96.55±0.19
96.94±0.23
LIBERO-10
92.40
95.23±0.14
93.44±0.06
Mean
96.85
97.09±0.04
96.97±0.07
Appendix
Table 19: Success rates (%) of π0.5 across LIBERO suites. Flow and HT report mean ± standard error across 10 evaluation seeds for each trained model. Bold marks the higher mean of Flow and HT.
ID
Task
Flow
HT
LIBERO-Spatial
0
Pick up the black bowl between the plate and the ramekin and place it on the plate.
99.90±0.10
99.90±0.10
1
Pick up the black bowl next to the ramekin and place it on the plate.
99.60±0.22
99.30±0.21
2
Pick up the black bowl from table center and place it on the plate.
99.70±0.21
99.70±0.15
3
Pick up the black bowl on the cookie box and place it on the plate.
99.40±0.22
98.80±0.36
4
Pick up the black bowl in the top drawer of the wooden cabinet and place it on the plate.
94.20±0.39
94.10±0.64
Appendix
Table 20: Task-level success rates (%) of π0.5 on LIBERO-Spatial and LIBERO-Object. Task scores are mean ± standard error over 10 evaluation seeds for each trained model. Bold marks the higher mean for each task.
ID
Task
Flow
HT
LIBERO-Goal
0
Open the middle drawer of the cabinet.
96.10±0.46
99.40±0.16
1
Put the bowl on the stove.
98.20±0.53
98.50±0.48
2
Put the wine bottle on top of the cabinet.
89.90±0.60
90.30±1.16
3
Open the top drawer and put the bowl inside.
92.50±0.48
95.40±0.85
4
Put the bowl on top of the cabinet.
99.40±0.16
99.50±0.27
Appendix
Table 21: Task-level success rates (%) of π0.5 on LIBERO-Goal and LIBERO-10. Task scores are mean ± standard error over 10 evaluation seeds for each trained model. Bold marks the higher mean for each task.
Task
Flow (official)
Flow
HT
PnPBottle → CabinetClose
70.00
42.00±4.23
58.26±5.53
PnPCan → DrawerClose
70.00
23.90±2.68
50.64±4.51
PnPCup → DrawerClose
35.00
8.40±1.83
23.58±3.32
PnPMilk → MicrowaveClose
45.00
11.00±2.56
25.00±4.22
PnPPotato → MicrowaveClose
40.00
6.00±1.25
26.31±3.84
PnPWine → CabinetClose
65.00
24.88±2.98
49.74±1.60
Appendix
Table 22: Task-level success rates (%) on RoboCasa-GR1. Flow and HT report mean ± standard error across 10 evaluation seeds for each trained model. Bold marks the higher mean of Flow and HT.
Task
Flow (official)
Flow
HT
carrot on plate
58.00
63.00±1.94
29.80±1.17
close drawer
97.00
98.40±0.50
98.80±0.53
put eggplant in basket
53.00
50.80±2.33
93.03±1.24
put eggplant in sink
2.00
2.00±0.84
65.00±1.82
open drawer
100.00
100.00±0.00
72.60±1.19
spoon on towel
78.00
87.02±1.13
61.40±1.98
Appendix
Table 23: Task-level success rates (%) on SIMPLER / Bridge. Flow and HT: mean ± standard error over 10 evaluation seeds. Mean averages seven tasks; bold marks the higher mean of Flow and HT.
Task
Flow (official)
Flow
HT
pick coke can
100.00
97.90±0.41
97.80±0.57
pick object
94.00
90.11±0.98
72.30±1.01
move near
100.00
98.50±0.50
93.10±0.87
open drawer
65.00
52.10±1.21
46.10±1.33
close drawer
69.00
58.84±0.97
66.20±1.19
place in closed drawer
7.00
9.50±1.01
16.00±0.54
Appendix
Table 24: Task-level success rates (%) on SIMPLER / Fractal. Flow and HT: mean ± standard error over 10 evaluation seeds. Mean averages six tasks; bold marks the higher mean of Flow and HT.
ID
Task
Flow
HT
0
Put both the alphabet soup and the tomato sauce in the basket.
92.00±4.16
95.00±2.24
1
Put both the cream cheese box and the butter in the basket.
100.00±0.00
100.00±0.00
2
Turn on the stove and put the moka pot on it.
94.00±2.21
99.00±1.00
3
Put the black bowl in the bottom drawer of the cabinet and close it.
94.00±2.21
97.00±1.53
4
Put the white mug on the left plate and put the yellow and white mug on the right plate.
95.00±1.67
96.00±2.21
5
Pick up the book and place it in the back compartment of the caddy.
99.00±1.00
98.00±1.33
Appendix
Table 25: Task-level success rates (%) of Cosmos3-Nano on LIBERO-10. Mean ± standard error over 10 evaluation seeds; bold marks the higher mean per row.
Action head
Whole model
Model
Method
NFE ↓
Latency ↓
Latency ↓
FPS ↑
Speedup ↑
GR00T N1.7
Flow
4
89.2
139.1
7.19
1.94 ×
HT (ours)
1
24.3
71.7
13.95
π0.5
Flow
10
169.7
260.8
3.83
2.42 ×
HT (ours)
1
17.1
107.6
9.29
Cosmos3-Nano
Flow
30
1368.2
1377.5
0.73
16.74 ×
Appendix
Table 26: Inference efficiency on an NVIDIA A800. Batch size one. All latency values are reported in milliseconds. NFE denotes the number of action-head evaluations per action chunk. Speedup is the Flow-to-HT whole-model latency ratio.
Objective
SR (%) ↑
MSE
37.80
L1
27.00
RMSE
46.00
Huber Huber (1964)
38.10
Student- t
39.50
Heteroscedastic Gaussian
42.50
Appendix
Table 27: Regression loss ablation. GR00T N1.7 on RoboCasa-GR1, 60k steps, 24 tasks. Upper/lower blocks omit/use input-dependent scales. Both Student- t variants use ν=928 .
Family
Fixed-scale penalty
Learned-scale penalty
Gaussian
S
S/(2σ2)+dlogσ
L1
Q
Q/σ+dlogσ
Radial norm (RMSE)
R
R/σ+dlogσ
Huber
hδ(R)
hδ(R/σ)+dlogσ
Student- t
Tν(R/σ0)+dlogσ0
Tν(R/σ)+dlogσ
Appendix
Table 28: Implemented regression penalties. Penalties are shown before division by the number of valid action coordinates. Let S=∥r∥22 , R=∥r∥2 , Q=∥r∥1 , and Tν(z)=2ν+dlog(1+z2/ν) . The learned scale σ=σ(o) is scalar.
Setting
Updates
d
Reported ν
RoboCasa-GR1
60,000
232
2,116,464,928,1450,∞ (HG)
Bridge
20,000
56
56,112,224,350,448
π0.5 / LIBERO
30,000
350
700,1400,2187.5
Appendix
Table 29: Degrees-of-freedom settings for Figure 11 . Bold marks the choice ν=4d ( c=4 ), which gives the highest tested success in each setting. HG denotes the Gaussian limit.
By relying on independent couplings from uninformative Gaussian priors, standard diffusion and flow matching models are forced to learn complex, high-cost vector fields to reach the physical action space. Generative models excel at capturing multimodal behaviors for robotic Learning from Demonstration (LfD), but often suffer from high inference cost. This paper introduces Temporal Policy, a generative framework based on stochastic interpolants that formulates action generation as a temporally coupled transport problem. By initializing the generative flow at the robot's recent history, we explicitly couple past states to future action sequences. This data-dependent coupling reduces transport cost and produces straight vector fields. We validate Temporal Policy across visuomotor simulation benchmarks and on a physical Barrett WAM 2x 7DoF teleoperation platform. Our approach reduces transport costs by nearly an order of magnitude compared to noise-initialized baselines, achieving a 19.1 ms inference latency on a single NVIDIA RTX 4080. Crucially, these geometric and computational efficiencies are achieved while matching the success rates of state-of-the-art baselines. This simplified transport geometry bypasses the computational bottleneck of independent Gaussian priors, helping enable high-frequency, closed-loop control. The code is publicly available at https://github.com/dmiller12/TemporalPolicy.
Dylan Miller, Martin Jagersand
Department of Computing Science, University of Alberta, Canada
Generative policies have emerged as a promising paradigm for robot learning, combining expressive generative action modeling with scalable imitation learning from large demonstration corpora. However, heterogeneous demonstrations can induce suboptimal action chunks whose errors compound over time, eventually driving the robot into out-of-distribution states from which recovery is difficult. Action verification offers a test-time scaling strategy for mitigating this failure mode by sampling multiple candidate actions and using a verifier to select one for execution. Existing approaches, however, remain temporally myopic and costly to train, evaluating candidates from the current observation alone without accounting for trajectory continuity and often relying on large verifiers and additional expert demonstrations. In this paper, we introduce Temporal Verification (TeV), an efficient temporally aware action verification framework for flow-matching VLAs. TeV first learns a temporal token that summarizes recent observation--action history, enabling candidate chunks to be evaluated as continuations of the execution trajectory rather than as isolated predictions. Conditioned on this token, TeV constructs positive--negative pairs without additional expert demonstrations or preference annotations and trains an energy-based verifier contrastively to assign lower energy to higher-quality, trajectory-consistent action chunks. Beyond post-hoc ranking, TeV further uses the learned energy landscape to guide intermediate flow samples toward lower-energy regions, improving candidates before final selection. Extensive experiments in simulation and real-world settings demonstrate that TeV provides reliably ranks action candidates, improves task success rates, and produces smoother execution trajectories.
Haoxuan Wang, Wayne Wu, Yan Yan +1
University of Illinois Chicago · University of California, Los Angeles
Generative vision-language-action policies have advanced robot manipulation, but they often exhibit instability under noise, partial observability, and stochastic initial conditions. During extended rollouts, small velocity errors accumulate, degrading execution reliability. Existing diffusion and flow-based policies typically assume homoscedastic residuals and lack explicit uncertainty modeling within action generation, limiting robustness during iterative rollout. We propose SUREFlow, a state-space uncertainty-aware residual flow matching framework built on a Mamba backbone. The method jointly predicts action velocities and input-dependent residual uncertainty, enabling selective refinement of unreliable action dimensions without environment feedback while preserving computational efficiency. On LIBERO, SUREFlow achieves 92.5% average success rate (SR), outperforming the Mamba-based MaIL by 34.2%. On LIBERO-PRO, it attains around 49% SR using only 179M parameters, achieving performance comparable to large VLAs with 3-7B parameters. SUREFlow source code is available on: https://github.com/tanvirnwu/SUREFlow
Md Tanvir Islam, Sai Navaneet Peddapalli, Sangmoon Lee +1
School of Electronic and Electrical Engineering, Kyungpook National University, Daegu 41566, Republic of Korea