Flow-matching policies can represent complex robot behaviors but remain susceptible to local errors under distribution shift at deployment. Many reinforcement learning approaches to policy improvement require reward signals that are difficult to specify or obtain in real-world manipulation. We present PreferenceFlow, a framework for improving a pretrained flow policy at test time without environment rewards or updates to the base policy. Human intervention chunks are paired with robot chunks generated from the same initial conditioning state to train a preference model. During infer- ence, we adopt the QGF sampling update, replacing its value gradient with the preference gradient evaluated at an estimated clean action. A gradient-cap loss penalizes excessive gradients on intervention pairs, while a zero-gradient loss discourages guidance near actions from expert demonstrations. On four real-world precision insertion tasks with a Franka robot, Pref- erenceFlow achieves a mean success rate of 90.5%, compared with 69% for the frozen policy. Ablations support the roles of gradient regularization and correctly ordered preference labels in the evaluated settings. These results demonstrate the utility of human interventions as local preference supervision for guiding frozen generative robot policies.
Figures & tables
Fig. 1: Overview of PreferenceFlow. Human interventions provide action-chunk comparisons sharing a window-start state, while expert demonstrations provide anchors for zero-gradient regularization. Separate intervention and demonstration sub-batches train a state-conditioned preference model. Its action gradient guides the frozen π0.5 flow sampler at inference. Only the preference-model parameters are updated during adaptation.
Fig. 2: State-conditioned preference-model architecture. Three camera images, the language instruction, and robot-state information form the frozen π0.5 conditioning representation, used as keys and values. Candidate action tokens form the queries and cross-attend to the prefix tokens to produce the preference score Fϕ(s,a) . Its action gradient provides the guidance signal for flow sampling. The schematic omits robot-state inputs and the per-step score heads followed by mean aggregation.
Fig. 3: Real-world precision insertion tasks. Each row shows one task and progresses from left to right, from the initial approach to insertion.
Method
Ethernet
USB
SATA
Charger
Mean ↑
Δ vs. base (pp) ↑
Frozen π0.5
66
70
68
72
69.0
—
Intervention BC
82
84
88
84
84.5
+15.5
QGF
74
74
80
76
76.0
+7.0
FlowDAgger
78
80
76
80
78.5
+9.5
PreferenceFlow (ours)
90
92
92
88
90.5
+21.5
TABLE I: Task success rates (%) over 50 real-robot trials per task. Mean denotes the arithmetic mean across the four tasks. Δ denotes the absolute improvement over the frozen policy in percentage points (pp). Bold indicates the best observed result in each column.
Training objective
Mean success (%) ↑
Fixed guidance
Adjusted guidance
Full objective
90.5
90.5
Without gradient-cap loss
15.0
81.0
Without zero-gradient loss
5.0
74.5
Preference loss only
Unstable
69.0
TABLE II: Gradient-objective ablation: mean success (%) across four tasks. Fixed guidance uses the full model’s weight. For adjusted guidance, each ablated variant’s weight is progressively reduced until deployment is stable. Unstable indicates that stable deployment was not possible at the tested weight; no success rate is assigned. Bold indicates the best observed rates.
Architecture
Mean success (%) ↑
Cross-attention + temporal self-attention
77.0
Per-step cross-attention (ours)
90.5
TABLE III: Preference-model architecture ablation: mean success (%) across four tasks. Both architectures use state-conditioned cross-attention; the ablated variant additionally uses temporal self-attention among action tokens. Bold indicates the best observed rate.
Variant
Mean success (%) ↑
Preference reranking (5 samples)
92.0
PreferenceFlow
90.5
Randomized preference labels
63.5
TABLE IV: Score-based selection and negative controls: mean success (%) across four tasks. Bold indicates the best observed rate.