Generative actors are transforming offline reinforcement learning (RL) by enabling expressive policy classes that model complex action distributions. However, this expressiveness also exposes a key challenge in heterogeneous datasets: generative policies can reproduce unreliable action modes whose return distributions exhibit high variance, occasionally yielding high returns by chance but lacking consistency. Consequently, maximizing the expected Q-value alone is insufficient for identifying reliable actions. We propose VAN-Flow (Variance-Averse n-step Flow), a framework that promotes reliable actions in generative offline RL. VAN-Flow combines (i) a categorical distributional critic, (ii) a variance-averse expectation operator that smoothly reweights atom probabilities to favor actions with both high returns and low dispersion, and (iii) a flow-matching generative actor guided via rejection sampling. Unlike CVaR or mean-variance objectives, the operator redistributes probability mass over the categorical return distribution without hard truncation or auxiliary penalty terms. Across more than 40 tasks from D4RL and OGBench, VAN-Flow consistently outperforms strong baselines, with the largest gains in long-horizon and high-variance regimes where reliable action selection becomes critical.
Figures & tables
Figure 1 : Success rate on antmaze-large under two dataset variance regimes: (a) navigate (low-variance) and (b) explore (high-variance). In (b), the shaded bar denotes the low-variance reference, and the downward arrow indicates the performance drop.
Figure 2 : Example of variance-averse return with varying δ and return variance.
Gaussian-based
Flow-based
Proposed
Task
IQL
HIQL
ReBRAC
ReBRAC- n
FQL
BFN
FQL- n
BFN- n
QC
VAN-Flow
Standard tasks
antmaze-large-navigate
53
91
81
82
79
71
91
81
49
95
antmaze-teleport-navigate
23
42
7
35
26
40
22
41
40
57
scene-play
0
38
11
40
57
85
18
57
84
93
puzzle-3x3-play
0
12
55
10
100
98
98
58
100
100
Table 1 : Performance comparison on OGBench. Each entry averages five tasks per environment. Best results are in bold.
Pure Offline RL
n -step
Dist.
Dist.+ n
Proposed
Task
IQL
ReBRAC
FQL
Retrace( λ )
PQL
LEQ
QC
TD3BC+MS
PA-RL
D4PG
VAN-Flow
umaze
77.0 ±5.5
97.7 ±1.5
96.0 ±2.0
96.0 ±3.0
94.0 ±3.0
94.4 ±6.3
95.0 ±1.4
–
–
87.6 ±4.1
98.0 ±2.8
umaze-diverse
54.2 ±5.5
83.5 ±7.0
89.0 ±2.0
84.0 ±10.0
63.0 ±6.0
71.0 ±12.3
78.0 ±8.5
–
–
61.3 ±9.4
95.6 ±2.6
medium-play
65.7 ±11.7
89.5 ±3.3
78.0 ±7.0
79.0 ±10.0
81.0 ±6.0
58.8 ±33.0
72.0 ±2.4
70.0
88.0
89.1 ±3.0
90.7 ±3.1
medium-diverse
73.7 ±5.4
83.5 ±8.2
71.0 ±13.0
53.0 ±17.0
60.0 ±33.0
46.2 ±23.2
64.0 ±2.8
66.0
88.0
81.3 ±4.1
89.3 ±5.0
large-play
42.0 ±4.5
52.2 ±29.0
84.0 ±7.0
12.0 ±26.0
33.0 ±10.0
58.6 ±9.1
80.0 ±2.8
22.0
87.0
71.3 ±4.1
86.5 ±5.7
Table 2 : Offline performance on D4RL AntMaze tasks. Scores are normalized following the D4RL protocol. Entries marked “–” indicate results not reported in prior work. Best results are in bold.
Figure 3 : Effect of flow-based policies and categorical critics on humanoidmaze-giant-navigate .
Figure 4 : Effect of Q guidance on policy performance across multiple tasks. Bars report mean success rate and time efficiency.
Figure 5 : Sensitivity analysis with respect to key algorithmic parameters.
Method
Complexity
Step (ms)
Wall (h)
FQL
O(UKd2)
1.1
0.32
BFN
O(UMKd2)
2.1
0.60
QC
O(UMKd2)
3.0
0.83
VAN-Flow
O(UMKd2)
1.9
0.54
Table 4 : Runtime comparison on humanoidmaze-large-navigate .
Appendix figures & tables23 assets
Supplementary material from the paper’s appendix.
Appendix
Notation
Description
M=(S,A,P,r,γ)
Markov decision process
πθ(⋅∣s)
Flow-based policy (actor) with parameters θ
vθπ(τ,s,x)
Flow-matching velocity field (policy parameterization)
Qψ(s,a)
Scalar critic (used in standard 1-step formulations)
Zψ(s,a)
Categorical return distribution critic used in VAN-Flow
Z(s,a)
Return random variable (categorical distribution) for (s,a)
Appendix
Table 5 : Notation table.
Measurement
Value
Median realized gap ∣E−Esp∣ (fraction of the value range R )
0.37%
States in which E and Esp select the same candidate
96.9%
Kendall’s τ between the two candidate rankings
0.99
Appendix
Table 6 : Discretization gap between the implemented operator E and its spectral form Esp on trained critics.
Figure 6 : D4RL
Figure 7 : OGBench
Hyperparameter
Value
Optimization
Learning rate (critic) αZ
3×10−4
Learning rate (actor) απ
3×10−4
Optimizer
Adam [ 54 ]
Gradient steps
106 (OGBench), 5×105 (D4RL)
Minibatch size
256
Appendix
Table 7 : Hyperparameters for VAN-Flow.
Variant
Dist
VE
Flow
antmaze-teleport (8d)
antmaze-giant (8d)
humanoidmaze-large (21d)
VAN-Flow (Ours)
✓
✓
✓
71±7
58±4
87±6
w/o VE
✓
✗
✓
38±7
36±6
77±4
w/o Flow
✓
✓
✗
42±7
38±13
0±0
w/o VE & Flow
✓
✗
✗
36±9
32±5
0±0
Appendix
Table 8 : Component-wise ablation across action-dimension regimes. All variants use n -step returns. Scores denote success rate (%), averaged over five random seeds.
VAN-DDPM
VAN-Flow
Steps
Perf.
Time (ms)
Perf.
Time (ms)
1
0\raisebox{0.68887pt}{\scriptsize\pm 0}
2.0
0\raisebox{0.68887pt}{\scriptsize\pm 0}
1.6
3
0\raisebox{0.68887pt}{\scriptsize\pm 0}
3.8
\mathbf{96\raisebox{0.68887pt}{\scriptsize\pm 3}}
3.3
5
0\raisebox{0.68887pt}{\scriptsize\pm 0}
5.7
96\raisebox{0.68887pt}{\scriptsize\pm 2}
5.0
10
57\raisebox{0.68887pt}{\scriptsize\pm 17}
8.4
87\raisebox{0.68887pt}{\scriptsize\pm 6}
7.2
20
\mathbf{94\raisebox{0.68887pt}{\scriptsize\pm 4}}
11.9
76\raisebox{0.68887pt}{\scriptsize\pm 6}
10.2
Appendix
Table 9 : Comparison between flow matching (VAN-Flow) and diffusion (VAN-DDPM) policies on humanoidmaze-large-navigate . Performance denotes success rate (%); Time denotes cumulative inference time per action (ms). Bold entries indicate the first step count achieving > 90% success rate.
Figure 8 : Learning curve on the antmaze task in OGBench
Figure 9 : Learning curve on the humanoidmaze task in OGBench
Figure 10 : Learning curve on the puzzle task in OGBench
Figure 11 : Learning curve on the scene task in OGBench
Figure 12 : Learning curve on the antmaze task in D4RL
Env
Task
Gaussian-based
Flow-based
Proposed
IQL
ReBRAC
ReBRAC- n
HIQL
FQL
BFN
FQL- n
BFN- n
QC
VAN-Flow
antmaze-large-navigate
task1
48 ±9
91 ±10
64.0 ±25.8
93.0 ±3.0
80 ±8
93.0 ±1.4
94.2 ±2.8
83.1 ±1.4
49.2 ±9.9
94.6 ±4.2
task2
42 ±6
88 ±4
78.3 ±5.9
78.0 ±9.0
57 ±10
90.2 ±2.8
80.0 ±3.4
88.0 ±11.3
43.4 ±2.4
89.4 ±7.0
task3
72 ±7
51 ±18
92.8 ±2.3
96.0 ±2.0
93 ±3
89.1 ±7.0
97.5 ±1.4
74.5 ±5.6
71.8 ±2.8
98.0 ±2.0
task4
51 ±9
84 ±7
90.0 ±2.8
94.0 ±2.0
80 ±4
9.0 ±1.4
89.1 ±7.0
73.1 ±1.4
22.0 ±2.8
95.4 ±3.1
task5
54 ±22
90 ±2
90.8 ±2.7
94.0 ±3.0
83 ±4
72.0 ±5.6
94.7 ±5.6
85.1 ±1.4
61.6 ±1.4
98.0 ±2.0
Appendix
Table 10 : Task-wise performance comparison in OGBench (singletask setting).
Table 12 : Normalized returns on D4RL locomotion tasks. Best results are in bold, and results within one standard deviation of the best are also bolded.
Task
VAN-Flow
QC
Task
VAN-Flow
QC
antmaze-teleport-navigate
57 → 96
40 → 46
scene-play
93 → 100
84 → 99
antmaze-large-navigate
95 → 99
49 → 98
puzzle-3x3-play
100 → 100
100 → 100
Long-horizon tasks
Noisy tasks
humanoidmaze-large-navigate
81 → 98
16 → 16
antmaze-large-explore
84 → 98
58 → 96
humanoidmaze-giant-navigate
92 → 98
12 → 67
puzzle-3x3-noisy
97 → 100
100 → 100
antmaze-giant-navigate
68 → 98
2 → 69
Appendix
Table 13 : Offline-to-online fine-tuning performance across diverse long-horizon and noisy control tasks. Scores are reported before online fine-tuning in gray and after fine-tuning in bold. The offline scores are those reported in Tables 1 and 2 .
Task
E[Z]
CVaR
CPW
Norm
Wang
Entropic
Sharpe
Sortino
Mean–Var.
VE
teleport
38±7
54±16
37±5
36±9
40±10
44±16
25±11
30±8
46±8
71±7
explore
84±2
2±1
87±4
84±8
86±2
86±4
64±27
70±8
85±4
93±3
Appendix
Table 14 : Comparison with representative risk-sensitive objectives. Scores denote success rates averaged over five random seeds.
Figure 13 : Additional sensitivity results on the number of n .
Figure 14 : Additional sensitivity results on the number of rejection samples
Figure 15 : Additional sensitivity results on the number of flow steps.