QAMM: Adjoint MeanFlow Matching for Few-Step Offline Reinforcement Learning
Organizations: School of Data Science, Fudan University · ShanghaiTech University · Shanghai Innovation Institute
Abstract
Flow policies can model rich action distributions, but their iterative sampling limits decision speed. Adjoint matching uses the critic's action gradient to improve a flow policy without backpropagating through its sampling trajectory, yet its supervision is defined for instantaneous velocities. We propose QAMM, a method that turns the critic-derived adjoint signal into supervision for MeanFlow's average velocity. The resulting policy learns finite-interval transport directly and generates actions with few network evaluations. We derive the adjoint MeanFlow target, specify its gradient boundaries, and train it with an offline actor-critic. On ten HumanoidMaze tasks, QAMM produces effective two-call policies and achieves competitive performance against strong flow-policy baselines. These results show that adjoint-based Q optimization can be combined with average-velocity learning to obtain expressive offline policies with few-step action generation.
Figures & tables
| Dataset | Task | FQL | CGQL-L | DSRL | IFQL | QAM | QAM-E | TRQAM | QAMM |
|---|---|---|---|---|---|---|---|---|---|
| Medium | 1 | ||||||||
| 2 | |||||||||
| 3 | |||||||||
| 4 | |||||||||
| 5 | |||||||||
| Mean |
Appendix figures & tables2 assets
Supplementary material from the paper’s appendix.
Appendix
| Setting | Value |
|---|---|
| Behavior pretraining / offline updates | 300k / 1M |
| Batch size / critic ensemble size | 256 / 10 |
| Action horizon / discount | 1 / 0.999 |
| Policy hidden layers | |
| Actor learning rate | |
| Behavior and critic learning rate |