cs.LGSep 28, 2026

QAMM: Adjoint MeanFlow Matching for Few-Step Offline Reinforcement Learning

Authors: Yuehu Gong, Shutong Ding, Mokai Pan, Yimiao Zhou, Jiashu Hou, Ye Shi, Yanwei Fu

Organizations: School of Data Science, Fudan University · ShanghaiTech University · Shanghai Innovation Institute

Abstract

Flow policies can model rich action distributions, but their iterative sampling limits decision speed. Adjoint matching uses the critic's action gradient to improve a flow policy without backpropagating through its sampling trajectory, yet its supervision is defined for instantaneous velocities. We propose QAMM, a method that turns the critic-derived adjoint signal into supervision for MeanFlow's average velocity. The resulting policy learns finite-interval transport directly and generates actions with few network evaluations. We derive the adjoint MeanFlow target, specify its gradient boundaries, and train it with an offline actor-critic. On ten HumanoidMaze tasks, QAMM produces effective two-call policies and achieves competitive performance against strong flow-policy baselines. These results show that adjoint-based Q optimization can be combined with average-velocity learning to obtain expressive offline policies with few-step action generation.

Figures & tables

Appendix figures & tables2 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Aligning Flow Map Policies with Optimal Q-Guidance

    May 12, 2026Christos Ziakas, Alessandra Russo, Avishek Joey BoseFlow PoliciesGeometric Guidance

  2. Trust Region Q Adjoint Matching

    May 26, 2026Yonghoon Dong, Kyungmin Lee, Changyeon Kim +2Flow PoliciesQ-Learning

  3. Entropy-Regularized Adjoint Matching for Offline Reinforcement Learning

    May 7, 2026Abdelghani Ghanem, Mounir GhoghoEntropy Regularized Reinforcement LearningModel-Based Reinforcement Learning