cs.LGOct 6, 2026

LASER: Latent Space Adjoint Matching for Support-Constrained Entropy-Regularized Offline RL

Authors: Songyuan Zhang, Oswin So, Eric Yang Yu, Matthew Cleaveland, Peter Crowley-Dolen, Chuchu Fan

Organizations: MIT · MIT Lincoln Laboratory

Abstract

While offline reinforcement learning (RL) enables policy optimization from static datasets without costly online interaction, it remains bottlenecked by the risk of executing out-of-distribution (OOD) actions. Recent approaches mitigate this by learning a behavior-cloning policy through flow matching and then performing RL within its constrained latent space. However, naively optimizing the latent policy can easily cause the policy to collapse into a brittle mode or exploit sharp artifacts of the learned critic. In this work, we find that entropy regularization is essential in latent-space RL for addressing these challenges. We introduce LASER, a novel offline RL algorithm that applies latent-space adjoint matching to achieve entropy-regularized latent-space RL with expressive flow policies while avoiding backpropagation through time. Through comprehensive experiments on 40 challenging OGBench tasks with varying dataset qualities, we show that LASER achieves state-of-the-art performance. Notably, LASER uses fixed method-specific hyperparameters across all tasks and outperforms the evaluated baselines, including those with task- and dataset-specific tuning, which highlights the robust applicability of LASER. Project website: https://mit-realm.github.io/laser/.

Figures & tables

Appendix figures & tables22 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Latent Policy Steering through One-Step Flow Policies

    Mar 5, 2026Hokyun Im, Andrey Kolobov, Jianlong Fu +1Model-Based Reinforcement LearningFlow Policies

  2. Entropy-Regularized Adjoint Matching for Offline Reinforcement Learning

    May 7, 2026Abdelghani Ghanem, Mounir GhoghoEntropy Regularized Reinforcement LearningModel-Based Reinforcement Learning

  3. Towards Efficient and Expressive Offline RL via Flow-Anchored Noise-conditioned Q-Learning

    May 3, 2026Sungyoung Lee, Dohyeong Kim, Eshan Balachandar +2Flow PoliciesModel-Based Reinforcement Learning