Generative actors are transforming offline reinforcement learning (RL) by enabling expressive policy classes that model complex action distributions. However, this expressiveness also exposes a key challenge in heterogeneous datasets: generative policies can reproduce unreliable action modes whose return distributions exhibit high variance, occasionally yielding high returns by chance but lacking consistency. Consequently, maximizing the expected Q-value alone is insufficient for identifying reliable actions. We propose VAN-Flow (Variance-Averse n-step Flow), a framework that promotes reliable actions in generative offline RL. VAN-Flow combines (i) a categorical distributional critic, (ii) a variance-averse expectation operator that smoothly reweights atom probabilities to favor actions with both high returns and low dispersion, and (iii) a flow-matching generative actor guided via rejection sampling. Unlike CVaR or mean-variance objectives, the operator redistributes probability mass over the categorical return distribution without hard truncation or auxiliary penalty terms. Across more than 40 tasks from D4RL and OGBench, VAN-Flow consistently outperforms strong baselines, with the largest gains in long-horizon and high-variance regimes where reliable action selection becomes critical.
Figures & tables
Figure 1 : Success rate on antmaze-large under two dataset variance regimes: (a) navigate (low-variance) and (b) explore (high-variance). In (b), the shaded bar denotes the low-variance reference, and the downward arrow indicates the performance drop.
Figure 2 : Example of variance-averse return with varying δ and return variance.
Gaussian-based
Flow-based
Proposed
Task
IQL
HIQL
ReBRAC
ReBRAC- n
FQL
BFN
FQL- n
BFN- n
QC
VAN-Flow
Standard tasks
antmaze-large-navigate
53
91
81
82
79
71
91
81
49
95
antmaze-teleport-navigate
23
42
7
35
26
40
22
41
40
57
scene-play
0
38
11
40
57
85
18
57
84
93
puzzle-3x3-play
0
12
55
10
100
98
98
58
100
100
Table 1 : Performance comparison on OGBench. Each entry averages five tasks per environment. Best results are in bold.
Pure Offline RL
n -step
Dist.
Dist.+ n
Proposed
Task
IQL
ReBRAC
FQL
Retrace( λ )
PQL
LEQ
QC
TD3BC+MS
PA-RL
D4PG
VAN-Flow
umaze
77.0 ±5.5
97.7 ±1.5
96.0 ±2.0
96.0 ±3.0
94.0 ±3.0
94.4 ±6.3
95.0 ±1.4
–
–
87.6 ±4.1
98.0 ±2.8
umaze-diverse
54.2 ±5.5
83.5 ±7.0
89.0 ±2.0
84.0 ±10.0
63.0 ±6.0
71.0 ±12.3
78.0 ±8.5
–
–
61.3 ±9.4
95.6 ±2.6
medium-play
65.7 ±11.7
89.5 ±3.3
78.0 ±7.0
79.0 ±10.0
81.0 ±6.0
58.8 ±33.0
72.0 ±2.4
70.0
88.0
89.1 ±3.0
90.7 ±3.1
medium-diverse
73.7 ±5.4
83.5 ±8.2
71.0 ±13.0
53.0 ±17.0
60.0 ±33.0
46.2 ±23.2
64.0 ±2.8
66.0
88.0
81.3 ±4.1
89.3 ±5.0
large-play
42.0 ±4.5
52.2 ±29.0
84.0 ±7.0
12.0 ±26.0
33.0 ±10.0
58.6 ±9.1
80.0 ±2.8
22.0
87.0
71.3 ±4.1
86.5 ±5.7
Table 2 : Offline performance on D4RL AntMaze tasks. Scores are normalized following the D4RL protocol. Entries marked “–” indicate results not reported in prior work. Best results are in bold.
Figure 3 : Effect of flow-based policies and categorical critics on humanoidmaze-giant-navigate .
Figure 4 : Effect of Q guidance on policy performance across multiple tasks. Bars report mean success rate and time efficiency.
Figure 5 : Sensitivity analysis with respect to key algorithmic parameters.
Method
Complexity
Step (ms)
Wall (h)
FQL
O(UKd2)
1.1
0.32
BFN
O(UMKd2)
2.1
0.60
QC
O(UMKd2)
3.0
0.83
VAN-Flow
O(UMKd2)
1.9
0.54
Table 4 : Runtime comparison on humanoidmaze-large-navigate .
Appendix figures & tables23 assets
Supplementary material from the paper’s appendix.
Appendix
Notation
Description
M=(S,A,P,r,γ)
Markov decision process
πθ(⋅∣s)
Flow-based policy (actor) with parameters θ
vθπ(τ,s,x)
Flow-matching velocity field (policy parameterization)
Qψ(s,a)
Scalar critic (used in standard 1-step formulations)
Zψ(s,a)
Categorical return distribution critic used in VAN-Flow
Z(s,a)
Return random variable (categorical distribution) for (s,a)
Appendix
Table 5 : Notation table.
Measurement
Value
Median realized gap ∣E−Esp∣ (fraction of the value range R )
0.37%
States in which E and Esp select the same candidate
96.9%
Kendall’s τ between the two candidate rankings
0.99
Appendix
Table 6 : Discretization gap between the implemented operator E and its spectral form Esp on trained critics.
Figure 6 : D4RL
Figure 7 : OGBench
Hyperparameter
Value
Optimization
Learning rate (critic) αZ
3×10−4
Learning rate (actor) απ
3×10−4
Optimizer
Adam [ 54 ]
Gradient steps
106 (OGBench), 5×105 (D4RL)
Minibatch size
256
Appendix
Table 7 : Hyperparameters for VAN-Flow.
Variant
Dist
VE
Flow
antmaze-teleport (8d)
antmaze-giant (8d)
humanoidmaze-large (21d)
VAN-Flow (Ours)
✓
✓
✓
71±7
58±4
87±6
w/o VE
✓
✗
✓
38±7
36±6
77±4
w/o Flow
✓
✓
✗
42±7
38±13
0±0
w/o VE & Flow
✓
✗
✗
36±9
32±5
0±0
Appendix
Table 8 : Component-wise ablation across action-dimension regimes. All variants use n -step returns. Scores denote success rate (%), averaged over five random seeds.
VAN-DDPM
VAN-Flow
Steps
Perf.
Time (ms)
Perf.
Time (ms)
1
0\raisebox{0.68887pt}{\scriptsize\pm 0}
2.0
0\raisebox{0.68887pt}{\scriptsize\pm 0}
1.6
3
0\raisebox{0.68887pt}{\scriptsize\pm 0}
3.8
\mathbf{96\raisebox{0.68887pt}{\scriptsize\pm 3}}
3.3
5
0\raisebox{0.68887pt}{\scriptsize\pm 0}
5.7
96\raisebox{0.68887pt}{\scriptsize\pm 2}
5.0
10
57\raisebox{0.68887pt}{\scriptsize\pm 17}
8.4
87\raisebox{0.68887pt}{\scriptsize\pm 6}
7.2
20
\mathbf{94\raisebox{0.68887pt}{\scriptsize\pm 4}}
11.9
76\raisebox{0.68887pt}{\scriptsize\pm 6}
10.2
Appendix
Table 9 : Comparison between flow matching (VAN-Flow) and diffusion (VAN-DDPM) policies on humanoidmaze-large-navigate . Performance denotes success rate (%); Time denotes cumulative inference time per action (ms). Bold entries indicate the first step count achieving > 90% success rate.
Figure 8 : Learning curve on the antmaze task in OGBench
Figure 9 : Learning curve on the humanoidmaze task in OGBench
Figure 10 : Learning curve on the puzzle task in OGBench
Figure 11 : Learning curve on the scene task in OGBench
Figure 12 : Learning curve on the antmaze task in D4RL
Env
Task
Gaussian-based
Flow-based
Proposed
IQL
ReBRAC
ReBRAC- n
HIQL
FQL
BFN
FQL- n
BFN- n
QC
VAN-Flow
antmaze-large-navigate
task1
48 ±9
91 ±10
64.0 ±25.8
93.0 ±3.0
80 ±8
93.0 ±1.4
94.2 ±2.8
83.1 ±1.4
49.2 ±9.9
94.6 ±4.2
task2
42 ±6
88 ±4
78.3 ±5.9
78.0 ±9.0
57 ±10
90.2 ±2.8
80.0 ±3.4
88.0 ±11.3
43.4 ±2.4
89.4 ±7.0
task3
72 ±7
51 ±18
92.8 ±2.3
96.0 ±2.0
93 ±3
89.1 ±7.0
97.5 ±1.4
74.5 ±5.6
71.8 ±2.8
98.0 ±2.0
task4
51 ±9
84 ±7
90.0 ±2.8
94.0 ±2.0
80 ±4
9.0 ±1.4
89.1 ±7.0
73.1 ±1.4
22.0 ±2.8
95.4 ±3.1
task5
54 ±22
90 ±2
90.8 ±2.7
94.0 ±3.0
83 ±4
72.0 ±5.6
94.7 ±5.6
85.1 ±1.4
61.6 ±1.4
98.0 ±2.0
Appendix
Table 10 : Task-wise performance comparison in OGBench (singletask setting).
Table 12 : Normalized returns on D4RL locomotion tasks. Best results are in bold, and results within one standard deviation of the best are also bolded.
Task
VAN-Flow
QC
Task
VAN-Flow
QC
antmaze-teleport-navigate
57 → 96
40 → 46
scene-play
93 → 100
84 → 99
antmaze-large-navigate
95 → 99
49 → 98
puzzle-3x3-play
100 → 100
100 → 100
Long-horizon tasks
Noisy tasks
humanoidmaze-large-navigate
81 → 98
16 → 16
antmaze-large-explore
84 → 98
58 → 96
humanoidmaze-giant-navigate
92 → 98
12 → 67
puzzle-3x3-noisy
97 → 100
100 → 100
antmaze-giant-navigate
68 → 98
2 → 69
Appendix
Table 13 : Offline-to-online fine-tuning performance across diverse long-horizon and noisy control tasks. Scores are reported before online fine-tuning in gray and after fine-tuning in bold. The offline scores are those reported in Tables 1 and 2 .
Task
E[Z]
CVaR
CPW
Norm
Wang
Entropic
Sharpe
Sortino
Mean–Var.
VE
teleport
38±7
54±16
37±5
36±9
40±10
44±16
25±11
30±8
46±8
71±7
explore
84±2
2±1
87±4
84±8
86±2
86±4
64±27
70±8
85±4
93±3
Appendix
Table 14 : Comparison with representative risk-sensitive objectives. Scores denote success rates averaged over five random seeds.
Figure 13 : Additional sensitivity results on the number of n .
Figure 14 : Additional sensitivity results on the number of rejection samples
Figure 15 : Additional sensitivity results on the number of flow steps.
We propose Flow-Anchored Noise-conditioned Q-Learning (FAN), a highly efficient and high-performing offline reinforcement learning (RL) algorithm. Recent work has shown that expressive flow policies and distributional critics improve offline RL performance, but at a high computational cost. Specifically, flow policies require iterative sampling to produce a single action, and distributional critics require computation over multiple samples (e.g., quantiles) to estimate value. To address these inefficiencies while maintaining high performance, we introduce FAN. Our method employs a behavior regularization technique that uses a single flow policy iteration and requires a single Gaussian noise sample for distributional critics. Our theoretical analysis of convergence and performance bounds demonstrates that these simplifications not only improve efficiency but also lead to superior task performance. Experiments on robotic manipulation and locomotion tasks demonstrate that FAN achieves state-of-the-art performance while significantly reducing both training and inference runtimes. We release our code at https://github.com/brianlsy98/FAN.
Sungyoung Lee, Dohyeong Kim, Eshan Balachandar +2
The University of Texas at Austin, Austin, TX, USA · Independent Researcher, Seoul, South Korea.
Iterative generative modeling techniques, such as flow matching, provide powerful tools to model complex behaviors for effective offline reinforcement learning (RL). In this work, we propose a new off-policy RL algorithm that trains a flow policy based on prior data. Our idea starts from the "expanded" Markov decision process (MDP) framework, which treats individual flow refinement steps as separate actions in an MDP. To enable off-policy RL within this framework, we apply two techniques: we generate virtual on-policy trajectories (by "reversing" flows) to make this framework compatible with prior data, and we apply a bias-and-variance reduction technique to mitigate the curse of horizon in off-policy RL. We call the resulting algorithm Reversal Q-learning (RQL). RQL has several advantages over previous flow-based RL methods: it does not suffer from backpropagation through time, makes better use of the learned value function, and directly trains the full, expressive flow policy. Through our experiments on 50 challenging simulated robotic tasks, we show that RQL leads to the best average offline RL performance compared to state-of-the-art flow-based offline RL algorithms.
Many reinforcement learning (RL) tasks have discrete action spaces, but most generative policy methods based on diffusion and flow matching are designed for continuous control. Meanwhile, generative policies usually rely heavily on offline datasets and offline-to-online RL is itself challenging, as the policy must improve from new interaction without losing useful behavior learned from static data. To address those challenges, we introduce DRIFT, an online fine-tuning method that updates an offline pretrained continuous-time Markov chain (CTMC) policy with an advantage-weighted discrete flow matching loss. To preserve useful pretrained knowledge, we add a path-space penalty that regularizes the full CTMC trajectory distribution, rather than only the final action distribution. For large discrete action spaces, we introduce a candidate-set approximation that updates the actor over a small subset of actions sampled from reference-policy rollouts and uniform exploration. Our theoretical analysis shows that the candidate-set error is controlled by missing target probability mass, and the induced CTMC generator error decreases as the candidate set covers more high-probability actions. Experiments on prevailing discrete action RL task show that our method provides stable offline-to-online improvement across all tasks, achieving the highest average score on Jericho with a simple GRU encoder while outperforming methods that use pretrained language models. Controlled experiments further confirm that the path-space penalty remains bounded during fine-tuning and that the CTMC generator adapts to shifted rewards faster than deterministic baselines. The candidate-set mechanism is supported by a stability analysis showing that the generator error decreases exponentially with candidate coverage.
Fairoz Nower Khan, Nabuat Zaman Nahim, Peizhong Ju
Department of Computer Science, University of Kentucky