cs.LGOct 6, 2026

Variance-Averse nn-Step Offline Reinforcement Learning for Sparse Long-Horizon Environments

Authors: Guhyeon Kang, Minhae Kwon

Organizations: Department of Electrical and Computer Engineering Sungkyunkwan University, Republic of Korea

Abstract

Generative actors are transforming offline reinforcement learning (RL) by enabling expressive policy classes that model complex action distributions. However, this expressiveness also exposes a key challenge in heterogeneous datasets: generative policies can reproduce unreliable action modes whose return distributions exhibit high variance, occasionally yielding high returns by chance but lacking consistency. Consequently, maximizing the expected QQ-value alone is insufficient for identifying reliable actions. We propose VAN-Flow (Variance-Averse nn-step Flow), a framework that promotes reliable actions in generative offline RL. VAN-Flow combines (i) a categorical distributional critic, (ii) a variance-averse expectation operator that smoothly reweights atom probabilities to favor actions with both high returns and low dispersion, and (iii) a flow-matching generative actor guided via rejection sampling. Unlike CVaR or mean-variance objectives, the operator redistributes probability mass over the categorical return distribution without hard truncation or auxiliary penalty terms. Across more than 40 tasks from D4RL and OGBench, VAN-Flow consistently outperforms strong baselines, with the largest gains in long-horizon and high-variance regimes where reliable action selection becomes critical.

Figures & tables

Appendix figures & tables23 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Towards Efficient and Expressive Offline RL via Flow-Anchored Noise-conditioned Q-Learning

    May 3, 2026Sungyoung Lee, Dohyeong Kim, Eshan Balachandar +2Flow PoliciesModel-Based Reinforcement Learning

  2. Reversal Q-Learning

    Jun 16, 2026Aditya Oberai, Seohong Park, Sergey LevineFlow PoliciesModel-Based Reinforcement Learning

  3. Discrete Flow Matching for Offline-to-Online Reinforcement Learning

    May 12, 2026Fairoz Nower Khan, Nabuat Zaman Nahim, Peizhong JuFlow PoliciesOffline Reinforcement Learning