While offline reinforcement learning (RL) enables policy optimization from static datasets without costly online interaction, it remains bottlenecked by the risk of executing out-of-distribution (OOD) actions. Recent approaches mitigate this by learning a behavior-cloning policy through flow matching and then performing RL within its constrained latent space. However, naively optimizing the latent policy can easily cause the policy to collapse into a brittle mode or exploit sharp artifacts of the learned critic. In this work, we find that entropy regularization is essential in latent-space RL for addressing these challenges. We introduce LASER, a novel offline RL algorithm that applies latent-space adjoint matching to achieve entropy-regularized latent-space RL with expressive flow policies while avoiding backpropagation through time. Through comprehensive experiments on 40 challenging OGBench tasks with varying dataset qualities, we show that LASER achieves state-of-the-art performance. Notably, LASER uses fixed method-specific hyperparameters across all tasks and outperforms the evaluated baselines, including those with task- and dataset-specific tuning, which highlights the robust applicability of LASER. Project website: https://mit-realm.github.io/laser/.
Figures & tables
Figure 1 : Both support constraints and entropy regularization are essential. Without support constraints, the policy generates OOD actions that have unreliable Q -values and thus fail. Without entropy regularization, the Q -function exhibits strange artifacts and the agents get stuck at the wall, resulting in failures on the left side. By combining support constraints and entropy regularization, LASER avoids OOD actions, regularizes the Q -function, and encourages policy exploration, resulting in a successful policy.
Figure 2 : Q -function and policy are poorly learned without entropy regularization. We plot the learned Q -functions (lighter colors denote higher value) at the state marked by the white circle in Figure 1 . Without entropy regularization, the policy becomes nearly deterministic and clusters at a false maximum which causes the agents to move up and get stuck at the wall ( Figure 1 ). Adding noise to the policy in TD learning (target policy smoothing [ 20 ] ) fixes the Q -function landscape, but the policy remains nearly deterministic and is suboptimal. With entropy regularization, the policy samples spread out more and yield a better regularized Q -function, enabling the agents to move around the wall and reach the target ( Figure 1 ).
Figure 3 : Flow latent policy + adjoint matching Pareto-dominates a squashed Gaussian latent policy. For every entropy and Q -value achieved by the Gaussian latent policy, there exists an operating point achievable by the flow latent policy that has both higher entropy and higher Q -value.
Figure 4 : LASER Algorithm. ( ) We train a flow-matching decoder {\mu_{{\theta_{\textsf{L}\mathrel{\mathchoice{\mkern 2.0mu\hbox{\hbox to3.45pt{\vbox to3.22pt{\pgfpicture\makeatletter\hbox{\hskip 0.0pt\lower 0.0pt\hbox to0.0pt{\lxSVG@begingroup@{_scopebegin=1} \lxSVG@begingroup@{stroke=#000000} \lxSVG@begingroup@{fill=#000000} \lxSVG@setlinewidth{\the\pgflinewidth}\lxSVG@begingroup@{stroke-width=0.4pt} \lx@inpgf@ignorespaces\nullfont\hbox to0.0pt{{}{}{}{}\lxSVG@discardpath\lxSVG@discardpath@clipped{M 0 0 L 0 4.45 L 4.77 4.45 L 4.77 0 Z} {{{}{}{{}}{} {{}{{\lx@inpgf@ignorespaces}}}{{}{\lx@inpgf@ignorespaces}}{}{{}{\lx@inpgf@ignorespaces}} {{{{\lx@inpgf@ignorespaces}}\lxSVG@begingroup@{_scopebegin=1} \lxSVG@transformcm{1.0}{0.0}{0.0}{1.0}{0.0pt}{0.0pt}\lxSVG@begingroup@{transform=matrix(1.0 0.0 0.0 1.0 0 0)} \pgfsys@hbox{58}\lxSVG@closescope }}}} {\lx@inpgf@ignorespaces}{\lx@inpgf@ignorespaces}{\lx@inpgf@ignorespaces}{\lx@inpgf@ignorespaces}\hss}\lxSVG@discardpath\lxSVG@closescope \hss}}\lxSVG@closescope\endpgfpicture}}}}{\mkern 2.0mu\hbox{\hbox to3.45pt{\vbox to3.22pt{\pgfpicture\makeatletter\hbox{\hskip 0.0pt\lower 0.0pt\hbox to0.0pt{\lxSVG@begingroup@{_scopebegin=1} \lxSVG@begingroup@{stroke=#000000} \lxSVG@begingroup@{fill=#000000} \lxSVG@setlinewidth{\the\pgflinewidth}\lxSVG@begingroup@{stroke-width=0.4pt} \lx@inpgf@ignorespaces\nullfont\hbox to0.0pt{{}{}{}{}\lxSVG@discardpath\lxSVG@discardpath@clipped{M 0 0 L 0 4.45 L 4.77 4.45 L 4.77 0 Z} {{{}{}{{}}{} {{}{{\lx@inpgf@ignorespaces}}}{{}{\lx@inpgf@ignorespaces}}{}{{}{\lx@inpgf@ignorespaces}} {{{{\lx@inpgf@ignorespaces}}\lxSVG@begingroup@{_scopebegin=1} \lxSVG@transformcm{1.0}{0.0}{0.0}{1.0}{0.0pt}{0.0pt}\lxSVG@begingroup@{transform=matrix(1.0 0.0 0.0 1.0 0 0)} \pgfsys@hbox{58}\lxSVG@closescope }}}} {\lx@inpgf@ignorespaces}{\lx@inpgf@ignorespaces}{\lx@inpgf@ignorespaces}{\lx@inpgf@ignorespaces}\hss}\lxSVG@discardpath\lxSVG@closescope \hss}}\lxSVG@closescope\endpgfpicture}}}}{\mkern 2.0mu\hbox{\hbox to2.42pt{\vbox to2.25pt{\pgfpicture\makeatletter\hbox{\hskip 0.0pt\lower 0.0pt\hbox to0.0pt{\lxSVG@begingroup@{_scopebegin=1} \lxSVG@begingroup@{stroke=#000000} \lxSVG@begingroup@{fill=#000000} \lxSVG@setlinewidth{\the\pgflinewidth}\lxSVG@begingroup@{stroke-width=0.4pt} \lx@inpgf@ignorespaces\nullfont\hbox to0.0pt{{}{}{}{}\lxSVG@discardpath\lxSVG@discardpath@clipped{M 0 0 L 0 3.11 L 3.34 3.11 L 3.34 0 Z} {{{}{}{{}}{} {{}{{\lx@inpgf@ignorespaces}}}{{}{\lx@inpgf@ignorespaces}}{}{{}{\lx@inpgf@ignorespaces}} {{{{\lx@inpgf@ignorespaces}}\lxSVG@begingroup@{_scopebegin=1} \lxSVG@transformcm{1.0}{0.0}{0.0}{1.0}{0.0pt}{0.0pt}\lxSVG@begingroup@{transform=matrix(1.0 0.0 0.0 1.0 0 0)} \pgfsys@hbox{58}\lxSVG@closescope }}}} {\lx@inpgf@ignorespaces}{\lx@inpgf@ignorespaces}{\lx@inpgf@ignorespaces}{\lx@inpgf@ignorespaces}\hss}\lxSVG@discardpath\lxSVG@closescope \hss}}\lxSVG@closescope\endpgfpicture}}}}{\mkern 2.0mu\hbox{\hbox to1.73pt{\vbox to1.61pt{\pgfpicture\makeatletter\hbox{\hskip 0.0pt\lower 0.0pt\hbox to0.0pt{\lxSVG@begingroup@{_scopebegin=1} \lxSVG@begingroup@{stroke=#000000} \lxSVG@begingroup@{fill=#000000} \lxSVG@setlinewidth{\the\pgflinewidth}\lxSVG@begingroup@{stroke-width=0.4pt} \lx@inpgf@ignorespaces\nullfont\hbox to0.0pt{{}{}{}{}\lxSVG@discardpath\lxSVG@discardpath@clipped{M 0 0 L 0 2.22 L 2.39 2.22 L 2.39 0 Z} {{{}{}{{}}{} {{}{{\lx@inpgf@ignorespaces}}}{{}{\lx@inpgf@ignorespaces}}{}{{}{\lx@inpgf@ignorespaces}} {{{{\lx@inpgf@ignorespaces}}\lxSVG@begingroup@{_scopebegin=1} \lxSVG@transformcm{1.0}{0.0}{0.0}{1.0}{0.0pt}{0.0pt}\lxSVG@begingroup@{transform=matrix(1.0 0.0 0.0 1.0 0 0)} \pgfsys@hbox{58}\lxSVG@closescope }}}} {\lx@inpgf@ignorespaces}{\lx@inpgf@ignorespaces}{\lx@inpgf@ignorespaces}{\lx@inpgf@ignorespaces}\hss}\lxSVG@discardpath\lxSVG@closescope \hss}}\lxSVG@closescope\endpgfpicture}}}}}\textsf{A}}}}} that maps samples from the uniform distribution in the d -ball qL=qLunif=U(Bld) in latent space to the dataset actions in action space . ( ) We then learn a base latent flow {\mu_{{\textsf{B}\mathrel{\mathchoice{\mkern 2.0mu\hbox{\hbox to5.95pt{\vbox to4.54pt{\pgfpicture\makeatletter\hbox{\hskip 0.0pt\lower 0.0pt\hbox to0.0pt{\lxSVG@begingroup@{_scopebegin=1} \lxSVG@begingroup@{stroke=#000000} \lxSVG@begingroup@{fill=#000000} \lxSVG@setlinewidth{\the\pgflinewidth}\lxSVG@begingroup@{stroke-width=0.4pt} \lx@inpgf@ignorespaces\nullfont\hbox to0.0pt{{}{}{}{}\lxSVG@discardpath\lxSVG@discardpath@clipped{M 0 0 L 0 6.28 L 8.23 6.28 L 8.23 0 Z} {{{}{}{{}}{} {{}{{\lx@inpgf@ignorespaces}}}{{}{\lx@inpgf@ignorespaces}}{}{{}{\lx@inpgf@ignorespaces}} {{{{\lx@inpgf@ignorespaces}}\lxSVG@begingroup@{_scopebegin=1} \lxSVG@transformcm{1.0}{0.0}{0.0}{1.0}{0.0pt}{0.0pt}\lxSVG@begingroup@{transform=matrix(1.0 0.0 0.0 1.0 0 0)} \pgfsys@hbox{58}\lxSVG@closescope }}}} {\lx@inpgf@ignorespaces}{\lx@inpgf@ignorespaces}{\lx@inpgf@ignorespaces}{\lx@inpgf@ignorespaces}\hss}\lxSVG@discardpath\lxSVG@closescope \hss}}\lxSVG@closescope\endpgfpicture}}}}{\mkern 2.0mu\hbox{\hbox to5.95pt{\vbox to4.54pt{\pgfpicture\makeatletter\hbox{\hskip 0.0pt\lower 0.0pt\hbox to0.0pt{\lxSVG@begingroup@{_scopebegin=1} \lxSVG@begingroup@{stroke=#000000} \lxSVG@begingroup@{fill=#000000} \lxSVG@setlinewidth{\the\pgflinewidth}\lxSVG@begingroup@{stroke-width=0.4pt} \lx@inpgf@ignorespaces\nullfont\hbox to0.0pt{{}{}{}{}\lxSVG@discardpath\lxSVG@discardpath@clipped{M 0 0 L 0 6.28 L 8.23 6.28 L 8.23 0 Z} {{{}{}{{}}{} {{}{{\lx@inpgf@ignorespaces}}}{{}{\lx@inpgf@ignorespaces}}{}{{}{\lx@inpgf@ignorespaces}} {{{{\lx@inpgf@ignorespaces}}\lxSVG@begingroup@{_scopebegin=1} \lxSVG@transformcm{1.0}{0.0}{0.0}{1.0}{0.0pt}{0.0pt}\lxSVG@begingroup@{transform=matrix(1.0 0.0 0.0 1.0 0 0)} \pgfsys@hbox{58}\lxSVG@closescope }}}} {\lx@inpgf@ignorespaces}{\lx@inpgf@ignorespaces}{\lx@inpgf@ignorespaces}{\lx@inpgf@ignorespaces}\hss}\lxSVG@discardpath\lxSVG@closescope \hss}}\lxSVG@closescope\endpgfpicture}}}}{\mkern 2.0mu\hbox{\hbox to3.38pt{\vbox to3.15pt{\pgfpicture\makeatletter\hbox{\hskip 0.0pt\lower 0.0pt\hbox to0.0pt{\lxSVG@begingroup@{_scopebegin=1} \lxSVG@begingroup@{stroke=#000000} \lxSVG@begingroup@{fill=#000000} \lxSVG@setlinewidth{\the\pgflinewidth}\lxSVG@begingroup@{stroke-width=0.4pt} \lx@inpgf@ignorespaces\nullfont\hbox to0.0pt{{}{}{}{}\lxSVG@discardpath\lxSVG@discardpath@clipped{M 0 0 L 0 4.36 L 4.68 4.36 L 4.68 0 Z} {{{}{}{{}}{} {{}{{\lx@inpgf@ignorespaces}}}{{}{\lx@inpgf@ignorespaces}}{}{{}{\lx@inpgf@ignorespaces}} {{{{\lx@inpgf@ignorespaces}}\lxSVG@begingroup@{_scopebegin=1} \lxSVG@transformcm{1.0}{0.0}{0.0}{1.0}{0.0pt}{0.0pt}\lxSVG@begingroup@{transform=matrix(1.0 0.0 0.0 1.0 0 0)} \pgfsys@hbox{58}\lxSVG@closescope }}}} {\lx@inpgf@ignorespaces}{\lx@inpgf@ignorespaces}{\lx@inpgf@ignorespaces}{\lx@inpgf@ignorespaces}\hss}\lxSVG@discardpath\lxSVG@closescope \hss}}\lxSVG@closescope\endpgfpicture}}}}{\mkern 2.0mu\hbox{\hbox to2.42pt{\vbox to2.25pt{\pgfpicture\makeatletter\hbox{\hskip 0.0pt\lower 0.0pt\hbox to0.0pt{\lxSVG@begingroup@{_scopebegin=1} \lxSVG@begingroup@{stroke=#000000} \lxSVG@begingroup@{fill=#000000} \lxSVG@setlinewidth{\the\pgflinewidth}\lxSVG@begingroup@{stroke-width=0.4pt} \lx@inpgf@ignorespaces\nullfont\hbox to0.0pt{{}{}{}{}\lxSVG@discardpath\lxSVG@discardpath@clipped{M 0 0 L 0 3.11 L 3.34 3.11 L 3.34 0 Z} {{{}{}{{}}{} {{}{{\lx@inpgf@ignorespaces}}}{{}{\lx@inpgf@ignorespaces}}{}{{}{\lx@inpgf@ignorespaces}} {{{{\lx@inpgf@ignorespaces}}\lxSVG@begingroup@{_scopebegin=1} \lxSVG@transformcm{1.0}{0.0}{0.0}{1.0}{0.0pt}{0.0pt}\lxSVG@begingroup@{transform=matrix(1.0 0.0 0.0 1.0 0 0)} \pgfsys@hbox{58}\lxSVG@closescope }}}} {\lx@inpgf@ignorespaces}{\lx@inpgf@ignorespaces}{\lx@inpgf@ignorespaces}{\lx@inpgf@ignorespaces}\hss}\lxSVG@discardpath\lxSVG@closescope \hss}}\lxSVG@closescope\endpgfpicture}}}}}\textsf{L}}}^{\mathrm{base}}}{} that maps the prior distribution qB in the base space to the uniform distribution in the latent space qLunif . ( ) The base latent flow is then fine-tuned via adjoint matching, which results in a latent flow policy {\mu_{{\theta_{\textsf{B}\mathrel{\mathchoice{\mkern 2.0mu\hbox{\hbox to3.45pt{\vbox to3.22pt{\pgfpicture\makeatletter\hbox{\hskip 0.0pt\lower 0.0pt\hbox to0.0pt{\lxSVG@begingroup@{_scopebegin=1} \lxSVG@begingroup@{stroke=#000000} \lxSVG@begingroup@{fill=#000000} \lxSVG@setlinewidth{\the\pgflinewidth}\lxSVG@begingroup@{stroke-width=0.4pt} \lx@inpgf@ignorespaces\nullfont\hbox to0.0pt{{}{}{}{}\lxSVG@discardpath\lxSVG@discardpath@clipped{M 0 0 L 0 4.45 L 4.77 4.45 L 4.77 0 Z} {{{}{}{{}}{} {{}{{\lx@inpgf@ignorespaces}}}{{}{\lx@inpgf@ignorespaces}}{}{{}{\lx@inpgf@ignorespaces}} {{{{\lx@inpgf@ignorespaces}}\lxSVG@begingroup@{_scopebegin=1} \lxSVG@transformcm{1.0}{0.0}{0.0}{1.0}{0.0pt}{0.0pt}\lxSVG@begingroup@{transform=matrix(1.0 0.0 0.0 1.0 0 0)} \pgfsys@hbox{58}\lxSVG@closescope }}}} {\lx@inpgf@ignorespaces}{\lx@inpgf@ignorespaces}{\lx@inpgf@ignorespaces}{\lx@inpgf@ignorespaces}\hss}\lxSVG@discardpath\lxSVG@closescope \hss}}\lxSVG@closescope\endpgfpicture}}}}{\mkern 2.0mu\hbox{\hbox to3.45pt{\vbox to3.22pt{\pgfpicture\makeatletter\hbox{\hskip 0.0pt\lower 0.0pt\hbox to0.0pt{\lxSVG@begingroup@{_scopebegin=1} \lxSVG@begingroup@{stroke=#000000} \lxSVG@begingroup@{fill=#000000} \lxSVG@setlinewidth{\the\pgflinewidth}\lxSVG@begingroup@{stroke-width=0.4pt} \lx@inpgf@ignorespaces\nullfont\hbox to0.0pt{{}{}{}{}\lxSVG@discardpath\lxSVG@discardpath@clipped{M 0 0 L 0 4.45 L 4.77 4.45 L 4.77 0 Z} {{{}{}{{}}{} {{}{{\lx@inpgf@ignorespaces}}}{{}{\lx@inpgf@ignorespaces}}{}{{}{\lx@inpgf@ignorespaces}} {{{{\lx@inpgf@ignorespaces}}\lxSVG@begingroup@{_scopebegin=1} \lxSVG@transformcm{1.0}{0.0}{0.0}{1.0}{0.0pt}{0.0pt}\lxSVG@begingroup@{transform=matrix(1.0 0.0 0.0 1.0 0 0)} \pgfsys@hbox{58}\lxSVG@closescope }}}} {\lx@inpgf@ignorespaces}{\lx@inpgf@ignorespaces}{\lx@inpgf@ignorespaces}{\lx@inpgf@ignorespaces}\hss}\lxSVG@discardpath\lxSVG@closescope \hss}}\lxSVG@closescope\endpgfpicture}}}}{\mkern 2.0mu\hbox{\hbox to2.42pt{\vbox to2.25pt{\pgfpicture\makeatletter\hbox{\hskip 0.0pt\lower 0.0pt\hbox to0.0pt{\lxSVG@begingroup@{_scopebegin=1} \lxSVG@begingroup@{stroke=#000000} \lxSVG@begingroup@{fill=#000000} \lxSVG@setlinewidth{\the\pgflinewidth}\lxSVG@begingroup@{stroke-width=0.4pt} \lx@inpgf@ignorespaces\nullfont\hbox to0.0pt{{}{}{}{}\lxSVG@discardpath\lxSVG@discardpath@clipped{M 0 0 L 0 3.11 L 3.34 3.11 L 3.34 0 Z} {{{}{}{{}}{} {{}{{\lx@inpgf@ignorespaces}}}{{}{\lx@inpgf@ignorespaces}}{}{{}{\lx@inpgf@ignorespaces}} {{{{\lx@inpgf@ignorespaces}}\lxSVG@begingroup@{_scopebegin=1} \lxSVG@transformcm{1.0}{0.0}{0.0}{1.0}{0.0pt}{0.0pt}\lxSVG@begingroup@{transform=matrix(1.0 0.0 0.0 1.0 0 0)} \pgfsys@hbox{58}\lxSVG@closescope }}}} {\lx@inpgf@ignorespaces}{\lx@inpgf@ignorespaces}{\lx@inpgf@ignorespaces}{\lx@inpgf@ignorespaces}\hss}\lxSVG@discardpath\lxSVG@closescope \hss}}\lxSVG@closescope\endpgfpicture}}}}{\mkern 2.0mu\hbox{\hbox to1.73pt{\vbox to1.61pt{\pgfpicture\makeatletter\hbox{\hskip 0.0pt\lower 0.0pt\hbox to0.0pt{\lxSVG@begingroup@{_scopebegin=1} \lxSVG@begingroup@{stroke=#000000} \lxSVG@begingroup@{fill=#000000} \lxSVG@setlinewidth{\the\pgflinewidth}\lxSVG@begingroup@{stroke-width=0.4pt} \lx@inpgf@ignorespaces\nullfont\hbox to0.0pt{{}{}{}{}\lxSVG@discardpath\lxSVG@discardpath@clipped{M 0 0 L 0 2.22 L 2.39 2.22 L 2.39 0 Z} {{{}{}{{}}{} {{}{{\lx@inpgf@ignorespaces}}}{{}{\lx@inpgf@ignorespaces}}{}{{}{\lx@inpgf@ignorespaces}} {{{{\lx@inpgf@ignorespaces}}\lxSVG@begingroup@{_scopebegin=1} \lxSVG@transformcm{1.0}{0.0}{0.0}{1.0}{0.0pt}{0.0pt}\lxSVG@begingroup@{transform=matrix(1.0 0.0 0.0 1.0 0 0)} \pgfsys@hbox{58}\lxSVG@closescope }}}} {\lx@inpgf@ignorespaces}{\lx@inpgf@ignorespaces}{\lx@inpgf@ignorespaces}{\lx@inpgf@ignorespaces}\hss}\lxSVG@discardpath\lxSVG@closescope \hss}}\lxSVG@closescope\endpgfpicture}}}}}\textsf{L}}}}}{} that generates the final action distribution \pi_{\textsf{A}}{}=({\mu_{{\theta_{\textsf{L}\mathrel{\mathchoice{\mkern 2.0mu\hbox{\hbox to3.45pt{\vbox to3.22pt{\pgfpicture\makeatletter\hbox{\hskip 0.0pt\lower 0.0pt\hbox to0.0pt{\lxSVG@begingroup@{_scopebegin=1} \lxSVG@begingroup@{stroke=#000000} \lxSVG@begingroup@{fill=#000000} \lxSVG@setlinewidth{\the\pgflinewidth}\lxSVG@begingroup@{stroke-width=0.4pt} \lx@inpgf@ignorespaces\nullfont\hbox to0.0pt{{}{}{}{}\lxSVG@discardpath\lxSVG@discardpath@clipped{M 0 0 L 0 4.45 L 4.77 4.45 L 4.77 0 Z} {{{}{}{{}}{} {{}{{\lx@inpgf@ignorespaces}}}{{}{\lx@inpgf@ignorespaces}}{}{{}{\lx@inpgf@ignorespaces}} {{{{\lx@inpgf@ignorespaces}}\lxSVG@begingroup@{_scopebegin=1} \lxSVG@transformcm{1.0}{0.0}{0.0}{1.0}{0.0pt}{0.0pt}\lxSVG@begingroup@{transform=matrix(1.0 0.0 0.0 1.0 0 0)} \pgfsys@hbox{58}\lxSVG@closescope }}}} {\lx@inpgf@ignorespaces}{\lx@inpgf@ignorespaces}{\lx@inpgf@ignorespaces}{\lx@inpgf@ignorespaces}\hss}\lxSVG@discardpath\lxSVG@closescope \hss}}\lxSVG@closescope\endpgfpicture}}}}{\mkern 2.0mu\hbox{\hbox to3.45pt{\vbox to3.22pt{\pgfpicture\makeatletter\hbox{\hskip 0.0pt\lower 0.0pt\hbox to0.0pt{\lxSVG@begingroup@{_scopebegin=1} \lxSVG@begingroup@{stroke=#000000} \lxSVG@begingroup@{fill=#000000} \lxSVG@setlinewidth{\the\pgflinewidth}\lxSVG@begingroup@{stroke-width=0.4pt} \lx@inpgf@ignorespaces\nullfont\hbox to0.0pt{{}{}{}{}\lxSVG@discardpath\lxSVG@discardpath@clipped{M 0 0 L 0 4.45 L 4.77 4.45 L 4.77 0 Z} {{{}{}{{}}{} {{}{{\lx@inpgf@ignorespaces}}}{{}{\lx@inpgf@ignorespaces}}{}{{}{\lx@inpgf@ignorespaces}} {{{{\lx@inpgf@ignorespaces}}\lxSVG@begingroup@{_scopebegin=1} \lxSVG@transformcm{1.0}{0.0}{0.0}{1.0}{0.0pt}{0.0pt}\lxSVG@begingroup@{transform=matrix(1.0 0.0 0.0 1.0 0 0)} \pgfsys@hbox{58}\lxSVG@closescope }}}} {\lx@inpgf@ignorespaces}{\lx@inpgf@ignorespaces}{\lx@inpgf@ignorespaces}{\lx@inpgf@ignorespaces}\hss}\lxSVG@discardpath\lxSVG@closescope \hss}}\lxSVG@closescope\endpgfpicture}}}}{\mkern 2.0mu\hbox{\hbox to2.42pt{\vbox to2.25pt{\pgfpicture\makeatletter\hbox{\hskip 0.0pt\lower 0.0pt\hbox to0.0pt{\lxSVG@begingroup@{_scopebegin=1} \lxSVG@begingroup@{stroke=#000000} \lxSVG@begingroup@{fill=#000000} \lxSVG@setlinewidth{\the\pgflinewidth}\lxSVG@begingroup@{stroke-width=0.4pt} \lx@inpgf@ignorespaces\nullfont\hbox to0.0pt{{}{}{}{}\lxSVG@discardpath\lxSVG@discardpath@clipped{M 0 0 L 0 3.11 L 3.34 3.11 L 3.34 0 Z} {{{}{}{{}}{} {{}{{\lx@inpgf@ignorespaces}}}{{}{\lx@inpgf@ignorespaces}}{}{{}{\lx@inpgf@ignorespaces}} {{{{\lx@inpgf@ignorespaces}}\lxSVG@begingroup@{_scopebegin=1} \lxSVG@transformcm{1.0}{0.0}{0.0}{1.0}{0.0pt}{0.0pt}\lxSVG@begingroup@{transform=matrix(1.0 0.0 0.0 1.0 0 0)} \pgfsys@hbox{58}\lxSVG@closescope }}}} {\lx@inpgf@ignorespaces}{\lx@inpgf@ignorespaces}{\lx@inpgf@ignorespaces}{\lx@inpgf@ignorespaces}\hss}\lxSVG@discardpath\lxSVG@closescope \hss}}\lxSVG@closescope\endpgfpicture}}}}{\mkern 2.0mu\hbox{\hbox to1.73pt{\vbox to1.61pt{\pgfpicture\makeatletter\hbox{\hskip 0.0pt\lower 0.0pt\hbox to0.0pt{\lxSVG@begingroup@{_scopebegin=1} \lxSVG@begingroup@{stroke=#000000} \lxSVG@begingroup@{fill=#000000} \lxSVG@setlinewidth{\the\pgflinewidth}\lxSVG@begingroup@{stroke-width=0.4pt} \lx@inpgf@ignorespaces\nullfont\hbox to0.0pt{{}{}{}{}\lxSVG@discardpath\lxSVG@discardpath@clipped{M 0 0 L 0 2.22 L 2.39 2.22 L 2.39 0 Z} {{{}{}{{}}{} {{}{{\lx@inpgf@ignorespaces}}}{{}{\lx@inpgf@ignorespaces}}{}{{}{\lx@inpgf@ignorespaces}} {{{{\lx@inpgf@ignorespaces}}\lxSVG@begingroup@{_scopebegin=1} \lxSVG@transformcm{1.0}{0.0}{0.0}{1.0}{0.0pt}{0.0pt}\lxSVG@begingroup@{transform=matrix(1.0 0.0 0.0 1.0 0 0)} \pgfsys@hbox{58}\lxSVG@closescope }}}} {\lx@inpgf@ignorespaces}{\lx@inpgf@ignorespaces}{\lx@inpgf@ignorespaces}{\lx@inpgf@ignorespaces}\hss}\lxSVG@discardpath\lxSVG@closescope \hss}}\lxSVG@closescope\endpgfpicture}}}}}\textsf{A}}}}})_{\sharp}({\mu_{{\theta_{\textsf{B}\mathrel{\mathchoice{\mkern 2.0mu\hbox{\hbox to3.45pt{\vbox to3.22pt{\pgfpicture\makeatletter\hbox{\hskip 0.0pt\lower 0.0pt\hbox to0.0pt{\lxSVG@begingroup@{_scopebegin=1} \lxSVG@begingroup@{stroke=#000000} \lxSVG@begingroup@{fill=#000000} \lxSVG@setlinewidth{\the\pgflinewidth}\lxSVG@begingroup@{stroke-width=0.4pt} \lx@inpgf@ignorespaces\nullfont\hbox to0.0pt{{}{}{}{}\lxSVG@discardpath\lxSVG@discardpath@clipped{M 0 0 L 0 4.45 L 4.77 4.45 L 4.77 0 Z} {{{}{}{{}}{} {{}{{\lx@inpgf@ignorespaces}}}{{}{\lx@inpgf@ignorespaces}}{}{{}{\lx@inpgf@ignorespaces}} {{{{\lx@inpgf@ignorespaces}}\lxSVG@begingroup@{_scopebegin=1} \lxSVG@transformcm{1.0}{0.0}{0.0}{1.0}{0.0pt}{0.0pt}\lxSVG@begingroup@{transform=matrix(1.0 0.0 0.0 1.0 0 0)} \pgfsys@hbox{58}\lxSVG@closescope }}}} {\lx@inpgf@ignorespaces}{\lx@inpgf@ignorespaces}{\lx@inpgf@ignorespaces}{\lx@inpgf@ignorespaces}\hss}\lxSVG@discardpath\lxSVG@closescope \hss}}\lxSVG@closescope\endpgfpicture}}}}{\mkern 2.0mu\hbox{\hbox to3.45pt{\vbox to3.22pt{\pgfpicture\makeatletter\hbox{\hskip 0.0pt\lower 0.0pt\hbox to0.0pt{\lxSVG@begingroup@{_scopebegin=1} \lxSVG@begingroup@{stroke=#000000} \lxSVG@begingroup@{fill=#000000} \lxSVG@setlinewidth{\the\pgflinewidth}\lxSVG@begingroup@{stroke-width=0.4pt} \lx@inpgf@ignorespaces\nullfont\hbox to0.0pt{{}{}{}{}\lxSVG@discardpath\lxSVG@discardpath@clipped{M 0 0 L 0 4.45 L 4.77 4.45 L 4.77 0 Z} {{{}{}{{}}{} {{}{{\lx@inpgf@ignorespaces}}}{{}{\lx@inpgf@ignorespaces}}{}{{}{\lx@inpgf@ignorespaces}} {{{{\lx@inpgf@ignorespaces}}\lxSVG@begingroup@{_scopebegin=1} \lxSVG@transformcm{1.0}{0.0}{0.0}{1.0}{0.0pt}{0.0pt}\lxSVG@begingroup@{transform=matrix(1.0 0.0 0.0 1.0 0 0)} \pgfsys@hbox{58}\lxSVG@closescope }}}} {\lx@inpgf@ignorespaces}{\lx@inpgf@ignorespaces}{\lx@inpgf@ignorespaces}{\lx@inpgf@ignorespaces}\hss}\lxSVG@discardpath\lxSVG@closescope \hss}}\lxSVG@closescope\endpgfpicture}}}}{\mkern 2.0mu\hbox{\hbox to2.42pt{\vbox to2.25pt{\pgfpicture\makeatletter\hbox{\hskip 0.0pt\lower 0.0pt\hbox to0.0pt{\lxSVG@begingroup@{_scopebegin=1} \lxSVG@begingroup@{stroke=#000000} \lxSVG@begingroup@{fill=#000000} \lxSVG@setlinewidth{\the\pgflinewidth}\lxSVG@begingroup@{stroke-width=0.4pt} \lx@inpgf@ignorespaces\nullfont\hbox to0.0pt{{}{}{}{}\lxSVG@discardpath\lxSVG@discardpath@clipped{M 0 0 L 0 3.11 L 3.34 3.11 L 3.34 0 Z} {{{}{}{{}}{} {{}{{\lx@inpgf@ignorespaces}}}{{}{\lx@inpgf@ignorespaces}}{}{{}{\lx@inpgf@ignorespaces}} {{{{\lx@inpgf@ignorespaces}}\lxSVG@begingroup@{_scopebegin=1} \lxSVG@transformcm{1.0}{0.0}{0.0}{1.0}{0.0pt}{0.0pt}\lxSVG@begingroup@{transform=matrix(1.0 0.0 0.0 1.0 0 0)} \pgfsys@hbox{58}\lxSVG@closescope }}}} {\lx@inpgf@ignorespaces}{\lx@inpgf@ignorespaces}{\lx@inpgf@ignorespaces}{\lx@inpgf@ignorespaces}\hss}\lxSVG@discardpath\lxSVG@closescope \hss}}\lxSVG@closescope\endpgfpicture}}}}{\mkern 2.0mu\hbox{\hbox to1.73pt{\vbox to1.61pt{\pgfpicture\makeatletter\hbox{\hskip 0.0pt\lower 0.0pt\hbox to0.0pt{\lxSVG@begingroup@{_scopebegin=1} \lxSVG@begingroup@{stroke=#000000} \lxSVG@begingroup@{fill=#000000} \lxSVG@setlinewidth{\the\pgflinewidth}\lxSVG@begingroup@{stroke-width=0.4pt} \lx@inpgf@ignorespaces\nullfont\hbox to0.0pt{{}{}{}{}\lxSVG@discardpath\lxSVG@discardpath@clipped{M 0 0 L 0 2.22 L 2.39 2.22 L 2.39 0 Z} {{{}{}{{}}{} {{}{{\lx@inpgf@ignorespaces}}}{{}{\lx@inpgf@ignorespaces}}{}{{}{\lx@inpgf@ignorespaces}} {{{{\lx@inpgf@ignorespaces}}\lxSVG@begingroup@{_scopebegin=1} \lxSVG@transformcm{1.0}{0.0}{0.0}{1.0}{0.0pt}{0.0pt}\lxSVG@begingroup@{transform=matrix(1.0 0.0 0.0 1.0 0 0)} \pgfsys@hbox{58}\lxSVG@closescope }}}} {\lx@inpgf@ignorespaces}{\lx@inpgf@ignorespaces}{\lx@inpgf@ignorespaces}{\lx@inpgf@ignorespaces}\hss}\lxSVG@discardpath\lxSVG@closescope \hss}}\lxSVG@closescope\endpgfpicture}}}}}\textsf{L}}}}})_{\sharp}q_{\textsf{B}}{} .
Figure 5 : Performance profile plots for the main evaluation suite with different datasets. For a given success rate τ (x-axis), we plot the fraction of runs with a success rate ≥τ [ 38 ] . Shading shows pointwise 95% bootstrap confidence intervals, computed by resampling training seeds within each task 2000 times. LASER attains the strongest empirical performance profile on both datasets.
Figure 6 : Ablation studies in cube-double-play-singletask-task2-v0 . Moderate values of the inverse temperature 1/α∈[5,20] yield comparable results. Adjoint matching achieves better and more stable results than BPTT. Using a different qL yields comparable final success rates in this task, consistent with the proposed separation between entropy regularization and behavior regularization. Lines show mean success rates, and shading spans the minimum and maximum across runs.
Appendix figures & tables22 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 7 : ReFORM fails to maximize the Q -value because of the “static flow” issue. For a given behavior policy and a ground-truth Q -function, the generated latent-space distribution ({\mu_{{\theta_{\textsf{B}\mathrel{\mathchoice{\mkern 2.0mu\hbox{\hbox to3.45pt{\vbox to3.22pt{\pgfpicture\makeatletter\hbox{\hskip 0.0pt\lower 0.0pt\hbox to0.0pt{\lxSVG@begingroup@{_scopebegin=1} \lxSVG@begingroup@{stroke=#000000} \lxSVG@begingroup@{fill=#000000} \lxSVG@setlinewidth{\the\pgflinewidth}\lxSVG@begingroup@{stroke-width=0.4pt} \lx@inpgf@ignorespaces\nullfont\hbox to0.0pt{{}{}{}{}\lxSVG@discardpath\lxSVG@discardpath@clipped{M 0 0 L 0 4.45 L 4.77 4.45 L 4.77 0 Z} {{{}{}{{}}{} {{}{{\lx@inpgf@ignorespaces}}}{{}{\lx@inpgf@ignorespaces}}{}{{}{\lx@inpgf@ignorespaces}} {{{{\lx@inpgf@ignorespaces}}\lxSVG@begingroup@{_scopebegin=1} \lxSVG@transformcm{1.0}{0.0}{0.0}{1.0}{0.0pt}{0.0pt}\lxSVG@begingroup@{transform=matrix(1.0 0.0 0.0 1.0 0 0)} \pgfsys@hbox{58}\lxSVG@closescope }}}} {\lx@inpgf@ignorespaces}{\lx@inpgf@ignorespaces}{\lx@inpgf@ignorespaces}{\lx@inpgf@ignorespaces}\hss}\lxSVG@discardpath\lxSVG@closescope \hss}}\lxSVG@closescope\endpgfpicture}}}}{\mkern 2.0mu\hbox{\hbox to3.45pt{\vbox to3.22pt{\pgfpicture\makeatletter\hbox{\hskip 0.0pt\lower 0.0pt\hbox to0.0pt{\lxSVG@begingroup@{_scopebegin=1} \lxSVG@begingroup@{stroke=#000000} \lxSVG@begingroup@{fill=#000000} \lxSVG@setlinewidth{\the\pgflinewidth}\lxSVG@begingroup@{stroke-width=0.4pt} \lx@inpgf@ignorespaces\nullfont\hbox to0.0pt{{}{}{}{}\lxSVG@discardpath\lxSVG@discardpath@clipped{M 0 0 L 0 4.45 L 4.77 4.45 L 4.77 0 Z} {{{}{}{{}}{} {{}{{\lx@inpgf@ignorespaces}}}{{}{\lx@inpgf@ignorespaces}}{}{{}{\lx@inpgf@ignorespaces}} {{{{\lx@inpgf@ignorespaces}}\lxSVG@begingroup@{_scopebegin=1} \lxSVG@transformcm{1.0}{0.0}{0.0}{1.0}{0.0pt}{0.0pt}\lxSVG@begingroup@{transform=matrix(1.0 0.0 0.0 1.0 0 0)} \pgfsys@hbox{58}\lxSVG@closescope }}}} {\lx@inpgf@ignorespaces}{\lx@inpgf@ignorespaces}{\lx@inpgf@ignorespaces}{\lx@inpgf@ignorespaces}\hss}\lxSVG@discardpath\lxSVG@closescope \hss}}\lxSVG@closescope\endpgfpicture}}}}{\mkern 2.0mu\hbox{\hbox to2.42pt{\vbox to2.25pt{\pgfpicture\makeatletter\hbox{\hskip 0.0pt\lower 0.0pt\hbox to0.0pt{\lxSVG@begingroup@{_scopebegin=1} \lxSVG@begingroup@{stroke=#000000} \lxSVG@begingroup@{fill=#000000} \lxSVG@setlinewidth{\the\pgflinewidth}\lxSVG@begingroup@{stroke-width=0.4pt} \lx@inpgf@ignorespaces\nullfont\hbox to0.0pt{{}{}{}{}\lxSVG@discardpath\lxSVG@discardpath@clipped{M 0 0 L 0 3.11 L 3.34 3.11 L 3.34 0 Z} {{{}{}{{}}{} {{}{{\lx@inpgf@ignorespaces}}}{{}{\lx@inpgf@ignorespaces}}{}{{}{\lx@inpgf@ignorespaces}} {{{{\lx@inpgf@ignorespaces}}\lxSVG@begingroup@{_scopebegin=1} \lxSVG@transformcm{1.0}{0.0}{0.0}{1.0}{0.0pt}{0.0pt}\lxSVG@begingroup@{transform=matrix(1.0 0.0 0.0 1.0 0 0)} \pgfsys@hbox{58}\lxSVG@closescope }}}} {\lx@inpgf@ignorespaces}{\lx@inpgf@ignorespaces}{\lx@inpgf@ignorespaces}{\lx@inpgf@ignorespaces}\hss}\lxSVG@discardpath\lxSVG@closescope \hss}}\lxSVG@closescope\endpgfpicture}}}}{\mkern 2.0mu\hbox{\hbox to1.73pt{\vbox to1.61pt{\pgfpicture\makeatletter\hbox{\hskip 0.0pt\lower 0.0pt\hbox to0.0pt{\lxSVG@begingroup@{_scopebegin=1} \lxSVG@begingroup@{stroke=#000000} \lxSVG@begingroup@{fill=#000000} \lxSVG@setlinewidth{\the\pgflinewidth}\lxSVG@begingroup@{stroke-width=0.4pt} \lx@inpgf@ignorespaces\nullfont\hbox to0.0pt{{}{}{}{}\lxSVG@discardpath\lxSVG@discardpath@clipped{M 0 0 L 0 2.22 L 2.39 2.22 L 2.39 0 Z} {{{}{}{{}}{} {{}{{\lx@inpgf@ignorespaces}}}{{}{\lx@inpgf@ignorespaces}}{}{{}{\lx@inpgf@ignorespaces}} {{{{\lx@inpgf@ignorespaces}}\lxSVG@begingroup@{_scopebegin=1} \lxSVG@transformcm{1.0}{0.0}{0.0}{1.0}{0.0pt}{0.0pt}\lxSVG@begingroup@{transform=matrix(1.0 0.0 0.0 1.0 0 0)} \pgfsys@hbox{58}\lxSVG@closescope }}}} {\lx@inpgf@ignorespaces}{\lx@inpgf@ignorespaces}{\lx@inpgf@ignorespaces}{\lx@inpgf@ignorespaces}\hss}\lxSVG@discardpath\lxSVG@closescope \hss}}\lxSVG@closescope\endpgfpicture}}}}}\textsf{L}}}}^{s}})_{\sharp}q_{\textsf{B}}{} concentrates along radial directions aligned with the flow, rather than moving toward the boundary of Bld . This makes the flow static, resulting in suboptimal performance compared with LASER .
Algorithm 1 LASER
Hyperparameter
Value
Learning rate
0.0003
Optimizer
Adam [ 67 ]
Maximum gradient norm
1 ( QAM , QAM-E ), 10 (others)
Target network smoothing coefficient
0.005
Discount factor γ
0.995
MLP dimensions
[512,512,512,512]
Appendix
Table 1 : Fixed common hyperparameters. These hyperparameters are shared by LASER and all the baselines and are fixed in all environments.
Environment
Dataset
Action chunk length
Pessimistic backup weight ρ
antmaze-large-navigate
clean
1
0.5
antmaze-large-explore
noisy
1
0.5
cube-single-play
clean
5
0.5
cube-single-noisy
noisy
1
0
cube-double-play
clean
5
0.5
cube-double-noisy
noisy
1
0
Appendix
Table 2 : Environment-specific common hyperparameters. These hyperparameters are shared by LASER and all baselines, but vary by environment and dataset.
Environment
Dataset
FQL ( α )
DSRL ( σz )
QAM ( τ )
QAM-E ( τ,σa )
antmaze-large-navigate
clean
3.0
1.25
10.0
(1.0, 0.1)
antmaze-large-explore
noisy
0.3
1.25
10.0
(1.0, 0.1)
cube-single-play
clean
300
0.5
1.0
(1.0, 0.01)
cube-single-noisy
noisy
30
0.5
1.0
(3.0, 0.5)
cube-double-play
clean
300
1.5
1.0
(1.0, 0.01)
cube-double-noisy
noisy
30
1.5
1.0
(1.0, 0.01)
Appendix
Table 3 : Environment- and dataset-specific unique hyperparameters. These hyperparameters are specific to individual algorithms and are tuned by environment and dataset. The method-specific hyperparameters of LASER , ReFORM , and IFQL remain fixed across environments, while their common settings follow Table 2 .
Algorithm
FQL
QAM
QAM-E
IFQL
DSRL
ReFORM
LASER
it/s
220
163
125
210
165
138
123
Appendix
Table 4 : Training throughput comparison. We report the number of iterations per second (it/s) for LASER and all the baselines.
Task
Dataset
NoER
ReFORM
LASER
cube-single-play-singletask-task1-v0
clean
69±27
99±2
100±0
cube-single-play-singletask-task2-v0
clean
59±36
96±3
100±0
cube-single-play-singletask-task3-v0
clean
87±5
98±3
100±0
cube-single-play-singletask-task4-v0
clean
51±18
87±2
99±1
cube-single-play-singletask-task5-v0
clean
74±14
86±13
93±5
cube-single-noisy-singletask-task1-v0
noisy
76±6
100±0
100±0
Appendix
Table 5 : Entropy regularization ablation on 20 cube-single and cube-double tasks. Success rates (%) are reported as mean ± standard deviation across 3 seeds, with 50 evaluation runs per seed. Bold denotes results at or above 95% of the best mean in each setting.
Dataset
NoER
TPS
LASER ( 1/α )
1/α=3
1/α=5
1/α=10
1/α=20
1/α=30
clean
55±7
49±24
50±14
78±19
89±1
79±12
30±42
noisy
51±12
61±17
31±19
72±15
81±10
83±18
83±19
Appendix
Table 6 : Comparison of NoER , TPS , and LASER with different inverse temperatures on cube-double task 2. Success rates (%) are reported as mean ± standard deviation across 3 seeds, with 50 evaluation runs per seed. Bold denotes results at or above 95% of the best mean for each dataset.
QAM
QAM+Mix ( η=0.03 )
QAM+Mix ( η=0.1 )
LASER
83±11
38±10
9±6
89±1
Appendix
Table 7 : Uniform-mixture ablation in action space on cube-double task 2 with the clean dataset. Success rates (%) are reported as mean ± standard deviation across 3 seeds, with 50 evaluation runs per seed. Bold denotes results at or above 95% of the best mean.
Task
FQL
QAM
QAM-E
IFQL
DSRL
ReFORM
LASER
1
5±4
2±2
63±14
50±13
65±24
86±6
97±3
2
13±9
0±0
74±24
2±0
17±7
6±7
87±2
3
5±7
3±2
18±12
87±1
63±41
81±7
84±7
4
1±1
0±0
33±15
32±4
31±21
33±25
60±3
5
0±0
1±1
53±26
2±3
14±14
6±3
69±18
Appendix
Table 8 : Results on the five puzzle-4x4 tasks with the clean dataset (100M). Success rates (%) are reported as mean ± standard deviation across 3 seeds, with 50 evaluation runs per seed. Bold denotes results at or above 95% of the best mean for each task.
Dataset
Task 1
Task 2
Task 3
Task 4
Task 5
clean
98.2
99.0
98.9
97.6
98.0
noisy
97.4
97.7
97.9
97.5
97.9
Appendix
Table 9 : Boundary-hit rates (%) for ReFORM on cube-double . Results are averaged across 3 seeds, with 50 evaluation episodes per seed and 10 latent flow steps per policy call.
Figure 8 : Additional ablation studies. LASER demonstrates robustness to the choice of σ0 . Furthermore, applying terminal-gradient clipping improves training stability. Reducing the parameter-gradient cap from 10 to 1 yields a comparable stabilizing effect. Lines show mean success rates, and shading spans the minimum and maximum across runs.
Baseline
FQL
QAM
QAM-E
IFQL
DSRL
ReFORM
Mean gap
20.02
18.70
17.75
47.50
16.33
11.82
p -value
9.08×10−6
3.47×10−4
6.44×10−4
3.22×10−11
1.35×10−3
3.42×10−5
Appendix
Table 10 : Paired statistical comparisons on the 40 main evaluation settings. Mean gaps are LASER minus each baseline in percentage points. One-sided p -values are computed from unrounded success rates averaged over 3 seeds per setting, with 50 evaluation runs per seed.
Task
Dataset
FQL
QAM
QAM-E
IFQL
DSRL
ReFORM
LASER
antmaze-large-navigate-task1-v0
clean
64±45
96±3
93±2
79±13
90±3
91±2
95±1
antmaze-large-navigate-task2-v0
clean
31±43
60±42
94±3
65±5
81±5
78±2
89±2
antmaze-large-navigate-task3-v0
clean
71±12
85±3
97±1
85±9
94±3
74±6
92±3
antmaze-large-navigate-task4-v0
clean
33±41
93±1
91±1
31±43
75±14
69±4
85±1
antmaze-large-navigate-task5-v0
clean
88±6
95±1
93±4
83±2
91±4
83±1
91±1
antmaze-large-explore-task1-v0
noisy
97±1
0±0
0±0
50±36
0±0
91±4
97±2
Appendix
Table 11 : Full results on 40 OGBench tasks. Success rates (%) are reported as mean ± standard deviation across 3 seeds, with 50 evaluation episodes per seed. Bold denotes results at or above 95% of the best mean in each setting, following [ 18 ] . We omit the -singletask tags from the task names to save space.
Figure 9 : Training curves for the antmaze-large environment with the clean dataset.
Figure 10 : Training curves for the antmaze-large environment with the noisy dataset.
Figure 11 : Training curves for the cube-single environment with the clean dataset.
Figure 12 : Training curves for the cube-single environment with the noisy dataset.
Figure 13 : Training curves for the cube-double environment with the clean dataset.
Figure 14 : Training curves for the cube-double environment with the noisy dataset.
Figure 15 : Training curves for the scene environment with the clean dataset.
Figure 16 : Training curves for the scene environment with the noisy dataset.
Offline reinforcement learning (RL) allows robots to learn from offline datasets without risky exploration. Yet, offline RL's performance often hinges on a brittle trade-off between (1) return maximization, which can push policies outside the dataset support, and (2) behavioral constraints, which typically require sensitive hyperparameter tuning. Latent steering offers a structural way to stay within the dataset support during RL, but existing offline adaptations commonly approximate action values using latent-space critics learned via indirect distillation, which can lose information and hinder convergence. We propose Latent Policy Steering (LPS), which enables high-fidelity latent policy improvement by backpropagating original-action-space Q-gradients through a differentiable one-step MeanFlow policy to update a latent-action-space actor. By eliminating proxy latent critics, LPS allows an original-action-space critic to guide end-to-end latent-space optimization, while the one-step MeanFlow policy serves as a behavior-constrained generative prior. This decoupling yields a robust method that works out-of-the-box with minimal tuning. Across OGBench and real-world robotic tasks, LPS achieves state-of-the-art performance and consistently outperforms behavioral cloning and strong latent steering baselines.
Hokyun Im, Andrey Kolobov, Jianlong Fu +1
Department of Artificial Intelligence, Yonsei University · Microsoft Research
Integrating expressive generative policies, such as flow-matching models, into offline reinforcement learning (RL) allows agents to capture complex, multi-modal behaviors. While Q-learning with Adjoint Matching (QAM) stabilizes policy optimization via the continuous adjoint method, it remains inherently bound to the fixed behavior distribution. This dependence induces a \textit{popularity bias} that can suppress high-reward actions in low-density regions, and creates a \textit{support binding} that restricts off-manifold exploration. Existing workarounds, such as appending \textit{residual} Gaussian policies, often re-introduce the expressivity bottlenecks associated with unimodal distributions. In this work, we propose \textit{Maximum Entropy Adjoint Matching} (ME-AM), a unified framework that addresses these limitations within the continuous flow formulation. ME-AM incorporates two mechanisms: (1) a Mirror Descent entropy maximization objective that mitigates the popularity bias to facilitate the extraction of optimal policies from offline datasets, and (2) a \textit{Mixture Behavior Prior} that broadens the geometric support to encompass out-of-distribution high-reward regions. By exploring this extended geometry, ME-AM identifies robust actions while preserving the absolute continuity of the generative vector field. Empirically, ME-AM demonstrates competitive or superior performance compared to prior state-of-the-art (SOTA) methods across a diverse suite of sparse-reward continuous control environments.
Abdelghani Ghanem, Mounir Ghogho
College of Computing Mohammed VI Polytechnic University Rabat 11103, Morocco
We propose Flow-Anchored Noise-conditioned Q-Learning (FAN), a highly efficient and high-performing offline reinforcement learning (RL) algorithm. Recent work has shown that expressive flow policies and distributional critics improve offline RL performance, but at a high computational cost. Specifically, flow policies require iterative sampling to produce a single action, and distributional critics require computation over multiple samples (e.g., quantiles) to estimate value. To address these inefficiencies while maintaining high performance, we introduce FAN. Our method employs a behavior regularization technique that uses a single flow policy iteration and requires a single Gaussian noise sample for distributional critics. Our theoretical analysis of convergence and performance bounds demonstrates that these simplifications not only improve efficiency but also lead to superior task performance. Experiments on robotic manipulation and locomotion tasks demonstrate that FAN achieves state-of-the-art performance while significantly reducing both training and inference runtimes. We release our code at https://github.com/brianlsy98/FAN.
Sungyoung Lee, Dohyeong Kim, Eshan Balachandar +2
The University of Texas at Austin, Austin, TX, USA · Independent Researcher, Seoul, South Korea.