Flow-based Vision-Language-Action (VLA) policies are typically trained by behavior cloning and thus do not explicitly optimize long-term task return. Critic guidance steers generation toward higher-value actions, but existing methods differentiate the critic through a one-step surrogate of the sampler and back-propagate a critic ensemble at every flow step. In contrast, here we propose Adjoint Guidance Flow (AGF), which amortizes trajectory-aware critic guidance into a lightweight guidance network while preserving the pretrained VLA policy. Specifically, we formulate critic-guided flow generation as a deterministic optimal control problem, whose optimal guidance is a costate that carries the terminal critic gradient back through the remaining flow, and regress the guidance network onto this costate while keeping both the VLA and critic frozen. This design provides favorable memory and throughput scaling during training, and inference needs one guidance-network forward pass per step, without the critic ensemble, back-propagation, or adjoint computation. Across LIBERO, RoboCasa, and LIBERO-Pro, AGF consistently improves pretrained VLAs, remains competitive with critic-guidance and policy-fine-tuning baselines, and is the most robust method when a single guidance strength is deployed across tasks. Compared with QGF, AGF runs 3.6× faster per guidance step with 7.0× fewer parameters, with comparable and even better performance, showing that critic guidance can be trajectory-aware and lightweight.
Figures & tables
Figure 1: Overview of AGF. (a) Training: the terminal critic gradient is carried back through the frozen flow and regressed into gϕ . (b) Inference: QGF back-propagates a critic ensemble per step; AGF runs one forward pass of gϕ . (c) Deployment on a real robot, with no critic ensemble on board.
Paradigm
Inference-time control
Guidance signal
Guidance computation at inference
Trained component
Pointwise guidance
✓
∇atQπ(s,a^0∣t)
critic forward + backward
–
Adjoint matching
×
∇atQπ(s,a0)
none
policy
AGF (Ours)
✓
∇atQπ(s,a0)
guidance forward
guidance net
Table 1: Comparison of critic-guidance paradigms. Pointwise guidance: DPS ( Chung et al., 2023 ) , QGF ( Zhou et al., 2026 ) , QPILOTS ( Ruan et al., 2026 ) , GAF ( Yang et al., 2026 ) ; adjoint matching: AM ( Domingo-Enrich et al., 2025 ) , QAM ( Li and Levine, 2026 ) ; a^0∣t is the Tweedie estimate.
Figure 2: Guidance-network architectures. (a) The critic ensemble predicts Qπ(s,a0) from the observation, robot state, and clean action chunk. (b) Ours: gϕ reuses one frozen critic encoder and predicts the trajectory-wise adjoint gϕ(s,at,t) with a FiLM-conditioned MLP.
Figure 3: Critic analysis. (Left) Distribution of critic values by scenario outcome. (Right) Critic values over normalized scenario progress. Pooled across all 4 LIBERO suites (40 tasks, 1,200 episodes).
Benchmark / policy
Suite
Calibration
Base
QAM
Q-BoN
QDPS
QGF
AGF (M=1)
AGF (M=4)
LIBERO / SmolVLA
Goal
Suite-level
76.6
83.6
78.0
79.8
80.6
80.2
81.0
Task-level
82.4
82.2
81.4
83.2
83.2
Object
Suite-level
89.6
94.6
90.6
89.4
93.6
92.4
93.4
Task-level
92.0
91.0
95.6
95.0
95.4
Spatial
Suite-level
71.6
71.4
71.2
68.0
73.0
72.2
74.6
Task-level
75.4
73.2
77.6
74.8
75.6
Table 2: Success rate (%) on LIBERO (SmolVLA), RoboCasa ( π0.5 ), and LIBERO-Pro (MolmoAct2). Suite means over tasks, 50 episodes each; LIBERO-Pro is averaged over five visual conditions (Appendix C.1 ). Bold and underline mark the best and second-best inference-time method per row; QAM fine-tunes the policy and is excluded from the ranking.
Figure 4: Training and inference efficiency. (a) Peak training memory and throughput of AGF versus QAM as the batch size grows, on one RTX 4090. (b) Parameters and runtime per guidance step of AGF versus vectorized QGF (SmolVLA). AGF ⋆ : trainable part of gϕ only.
Figure 5: Weight ablation. Guidance-scale sweep ( 0.25 – 4.0 ) for each inference-time method, averaged over all LIBERO suites; AGF consistently improves performance across all scales.
Figure 6: Real-robot evaluation on open-pnp-close . Columns show top-view observations at identical timestamps across methods, each chosen where AGF completes a subtask. Insets show the right wrist camera; green and red borders mark subtask success and failure.
Base
QGF
AGF
pnp-plate ( n=50 )
78.0
78.0 (4/4)
88.0 (7/2)
open-pnp-close ( n=25 )
44.0
48.0 (5/4)
76.0 (10/2)
Pooled ( n=75 )
66.7
68.0 (9/8)
84.0 (17/4)
McNemar p (pooled)
–
1.000
0.007
Overhead (ms/chunk)
–
31.2
11.2
Table 3: Real robot results. Success rate (%) over paired episodes; parentheses count discordant pairs against Base (F → S, S → F), on which the pooled McNemar p is computed. Overhead is per action chunk.
Appendix figures & tables12 assets
Supplementary material from the paper’s appendix.
Appendix
Setting
Chunk length
Executed steps
Critic/guidance horizon
SmolVLA + LIBERO
50
10
10
π0.5 + RoboCasa
50
50
50
MolmoAct2 + LIBERO-Pro
10
10
10
MolmoAct2 + real robot †
30
∼ 15 (async)
29
Appendix
Table 4: Action-chunking configuration per setting. The critic and the guidance network share the same horizon. † For the details, see Appendix B.3 .
Figure 7: Real-robot setup. Bimanual platform with the top-view camera; the platform is used in a single-arm configuration and only one follower arm is controlled. Wrist camera on that arm.
Task
Episodes
Frames
Avg. Length (s)
pnp-plate †
106
47,002
14.8
pnp-box
108
47,616
14.7
open-pnp-close †
108
91,614
28.3
open-place
102
82,313
26.9
Overall
424
268,545
21.2
Appendix
Table 5: Real-world dataset statistics. † denotes the two evaluation tasks.
Figure 8: Paired evaluation interface. (A) Coverage tracker over the 50 reference episodes. (B) Reference view captured when the first method is evaluated, (C) the current scene before a subsequent method is run, and (D) their difference map, used to verify that each scenario is initialized consistently across methods.
Calibration
Method
Clean
Position
Object
Semantic
Environment
Avg.
Suite-level
Pretrained VLA
97.2
23.0
83.5
97.2
74.0
75.0
Q-BoN ( N=4 )
8.2
1.2
2.0
6.5
3.8
4.3
QDPS
97.5
24.0
84.0
98.2
76.0
76.0
QGF
97.2
24.2
85.2
98.0
75.8
76.1
AGF ( M=1 )
97.0
24.8
84.8
97.5
75.8
76.0
AGF ( M=4 )
98.0
25.8
84.0
97.8
75.2
76.2
Appendix
Table 6: Per-condition LIBERO-Pro results with MolmoAct2. Task success rate (%) averaged over the four LIBERO suites for each visual condition, under suite-level and task-level calibration; the Avg. column corresponds to the LIBERO-Pro rows of Table 2 . Bold and underline mark the best and second-best inference-time method per column within each calibration setting; our method is shaded.
N
Suite
4
8
16
Goal
Success rate
10.4
2.0
0.6
Qˉbest−Qˉmean
0.224
0.344
0.446
Object
Success rate
1.0
0.0
0.0
Qˉbest−Qˉmean
0.218
0.309
0.359
Spatial
Success rate
5.8
1.2
0.2
Appendix
Table 7: Best-of- N over-optimization on MolmoAct2. Suite-level success rate (%) and the critic’s apparent advantage of the selected candidate ( Qˉbest−Qˉmean , averaged over episodes) as the number of candidates N grows. Across all four suites, the apparent advantage increases monotonically while the actual success rate collapses.
LIBERO (SmolVLA)
RoboCasa ( π0.5 )
LIBERO-Pro (MolmoAct2)
Calibration
Method
Goal
Object
Spatial
LIBERO-10
Atomic
Goal
Object
Spatial
LIBERO-10
Task-level
QGF
81.4
95.6
77.6
39.2
49.6
77.4
85.8
77.0
69.4
AGF
83.2
95.4
75.6
42.0
49.4
77.8
84.6
79.0
68.8
Suite-level
QGF
80.6
93.6
73.0
36.0
46.3
76.0
84.6
76.0
67.8
AGF
81.0
93.4
74.6
36.8
46.6
76.4
83.8
77.4
67.0
LOTO
QGF
80.6
91.8
69.6
31.8
46.3
74.8
84.6
76.0
67.8
Appendix
Table 8: Leave-one-task-out (LOTO) calibration. Success rate (%) when the guidance strength is selected on all but one task and evaluated on the held-out task, averaged over held-out tasks. Bold marks the better LOTO result per suite; our method is shaded.
Comparison
Suite
Δ SR
McNemar p
95% CI
LIBERO (SmolVLA)
Base vs. AGF
Goal
+4.4
0.021
[+0.8,+8.0]
Object
+3.8
0.020
[+0.8,+6.8]
Spatial
+3.0
0.191
[−1.2,+7.2]
LIBERO-10
+3.2
0.195
[−1.2,+7.8]
Pooled
+3.6
3.4×10−4
[+1.7,+5.6]
Appendix
Table 9: Paired significance tests. Episode-level paired comparisons on identical seeds across all three benchmarks; positive Δ SR favors AGF.
LIBERO (SmolVLA)
RoboCasa ( π0.5 )
LIBERO-Pro (MolmoAct2)
Method
Goal
Object
Spatial
LIBERO-10
Atomic
Goal
Object
Spatial
LIBERO-10
QDPS
4.0
0.25
4.0
1.0
4.0
2.0
1.0
2.0
1.0
QGF
4.0
1.0
2.0
1.0
0.25
2.0
0.5
0.25
0.25
Q-BoN ( N )
8
8
4
4
4
4
4
4
4
AGF ( M=1 )
4.0
2.0
2.0
2.0
2.0
1.0
2.0
2.0
0.5
AGF ( M=4 )
4.0
2.0
4.0
4.0
4.0
2.0
4.0
1.0
2.0
Appendix
Table 10: Selected guidance hyperparameters. Suite-level selected guidance weight w for each gradient-based method and the selected number of candidates N for Q-BoN. All gradient-based methods share the sweep range [0.25,4.0] after flow-velocity normalization; Q-BoN sweeps N∈{4,8,16} . Ties are broken toward the smaller weight.
Subtask
Base
QGF
AGF
Open the box
15/25
18/25
25/25
Place the ball inside
12/25
15/25
21/25
Close the box
17/25
14/25
20/25
All three (task success)
11/25
12/25
19/25
Appendix
Table 11: Subtask success on open-pnp-close (25 paired episodes; each subtask judged independently, so a later subtask can succeed after an earlier one fails, e.g., closing the box without placing the ball).
Figure 9: Qualitative rollouts on pnp-plate . Eight timestamps from paired episodes of Base VLA, QGF, and AGF, all initialized identically. Within each block the top row is the wrist camera and the bottom row the top-view camera. Frames are dimmed after a method has completed the task.
Figure 10: Qualitative rollouts on open-pnp-close . Eight timestamps from paired episodes of Base VLA, QGF, and AGF, all initialized identically. Within each block the top row is the wrist camera and the bottom row the top-view camera. Frames are dimmed after a method has completed the task.
Brain-inspired Cognitive AI Lab, Institute of Automation, Chinese Academy of Sciences, Beijing, China. · Beijing Institute of AI Safety and Governance, China. · University of Chinese Academy of Sciences (UCAS), Beijing, China. +3