Q-Learning with Scalar Adjoint Matching
Organizations: KAIST · RLWRLD
Abstract
Flow policies capture rich and diverse action distributions, and fine-tuning them with off-policy RL to improve beyond the demonstrations has drawn growing interest. However, fine-tuning a flow policy against a learned value function is not trivial, because the policy generates its action over many flow steps. Adjoint matching offers a principled way to update the flow model itself by propagating value information from the final action back to each flow step, but it requires a vector--Jacobian product through the policy at every step, a cost that grows with the number of flow steps and the policy size. We observe that the batch-averaged velocity Jacobian of pretrained flow policies concentrates on its diagonal. Motivated by this finding, we derive a closed-form scalar adjoint that scales the value gradient at the final action by the flow time, eliminating the per-step vector--Jacobian products. We further find that controlling the critic's value at policy-generated actions is particularly important under the scalar adjoint. Based on these findings, we propose Q-learning with Scalar Adjoint Matching (SQAM), which combines the scalar adjoint with a value penalty at those actions. SQAM's gains concentrate on the four hardest OGBench domains, where its success rate exceeds that of the strongest baseline in each domain by 18 to 35 percentage points. To test whether SQAM extends to large pretrained policies, we also fine-tune a vision-language-action policy on a real bimanual robot. SQAM improves over supervised fine-tuning on all three tasks.
Figures & tables
| al | ag | hm | hl | scene | p33 | p44 | c2 | c3 | c4 | all | ||
| 5 tasks | 5 tasks | 5 tasks | 5 tasks | 5 tasks | 5 tasks | 5 tasks | 5 tasks | 5 tasks | 5 tasks | 50 tasks | ||
| Backprop | FQL | 38 9 | 2 6 | 74 5 | 2 1 | 70 5 | 25 10 | 9 7 | 44 4 | 7 5 | 9 5 | 28 |
| Guidance | CGQL-L | 48 7 | 7 5 | 57 2 | 6 3 | 58 1 | 0 0 | 0 0 | 55 2 | 0 1 | 1 1 | 23 |
| Post Processing | DSRL | 53 2 | 1 1 | 53 10 | 1 1 | 80 0 | 100 0 | 61 8 | 72 4 | 34 6 | 9 3 | 46 |
| IFQL | 29 8 | 12 3 | 93 2 | 30 7 | 36 1 | 64 4 | 42 4 | 9 2 | 24 7 | 6 3 | 35 | |
| Adjoint Matching | QAM | 62 9 | 29 4 | 64 7 | 4 3 | 64 4 | 15 3 | 1 1 | 71 2 | 19 6 | 18 3 | 35 |
| Method | Flip plastic bag | Place the straw | Put the fruit and close the lid |
| SFT | |||
| EXPO-FT (offline) | |||
| EXPO-FT (offline2online) | |||
| SQAM (offline) | |||
| SQAM (offline2online) |
Appendix figures & tables24 assets
Supplementary material from the paper’s appendix.
Appendix
| {\color[rgb]{0.2578,0.5234,0.957}X_{\tau+h}=X_{\tau}+h\,v^{\mathrm{ft}}_{\theta}(s,X_{\tau},\tau),\quad X_{0}\sim\mathcal{N}(0,I),} |
| Domain | Category | Data | Horizon | Act. dim. |
| cube-double-* | Manipulation | 1M | 500 | 5 |
| cube-triple-10M-* | Manipulation | 10M | 1000 | 5 |
| cube-quadruple-100M-* | Manipulation | 100M | 1000 | 5 |
| antmaze-large-* | Locomotion | 1M | 1000 | 8 |
| antmaze-giant-10M-* | Locomotion | 10M | 1000 | 8 |
| humanoidmaze-medium-* | Locomotion | 1M | 2000 | 21 |
| Parameter | Value |
| Batch size | 256 |
| Discount factor ( ) | 0.995 (default), 0.999 ( humanoidmaze ) |
| Optimizer | Adam |
| Learning rate | |
| Target network update rate | |
| Critic ensemble size ( ) | 10 |
| Domain | FQL | DSRL | IFQL | CGQL-L | QAM | QAM-E | TRQAM | SQAM |
| scene-* | 300 | 0.4 | 0.9 | 1 | 0.5 | |||
| puzzle-3x3-* | 300 | 1.0 | 0.95 | 3 | 2.0 | |||
| puzzle-4x4-10M-* | 1 | 1.0 | 0.9 | 30 | 4.0 | |||
| cube-double-* | 300 | 1.0 | 0.9 | 1 | 0.5 | |||
| cube-triple-10M-* | 30 | 1.4 | 0.95 | 3 | 0.5 |
| Task | Language instruction | Demos | |
| flip-plastic-bag | “flip the plastic bag” | 50 | 0.999 |
| place-straw | “place the straw in the cup” | 50 | 0.999 |
| put-and-close | “put the fruit in the pot and close the lid” | 50 | 0.999 |
| Parameter | Value |
| Shared | |
| Offline RL steps | 15,000 |
| Online RL steps | 10,000 |
| Batch size | 128 offline, 192 online |
| Policy delay | 5 |
| UTD ratio | 1 |