Wiring Matters: Injection Topology and Initialization of Affordance Heads in Vision-Language-Action Policies
Authors: Zijian An, Linhan Wang, Jiayan Wang, Shijie Geng, Ran Yang, Yiming Feng, Lifeng Zhou
Organizations: Drexel University, Philadelphia, PA 19104, USA · Virginia Tech, Virginia Seafood Agricultural Research and Extension Center, Hampton, VA 23669, USA · TODO: Jiayan Wang’s affiliation and email · Amazon Store Foundation AI (SFAI), New York, NY 10018, USA
Dense affordance supervision is an appealing auxiliary signal for vision-language-action (VLA) policies, yet naively co-training an affordance head can severely damage instruction following. We present a controlled study of how to wire such a head into a modern VLA on the LIBERO benchmark. Our recipe reads the backbone through a stop-gradient and re-injects an intermediate head feature into the action expert via a learned bridge. The stop-gradient is a precondition: letting affordance gradients reach the backbone drops the policy below the headless base (85.5% vs. 93.1%). With the backbone protected, a same-budget 2*2 ablation over injection topology (concatenation vs. residual) and bridge initialization (zero vs. random) shows initialization is the dominant lever. The best wiring, an actively initialized residual bridge, reaches 96.2%, matching the far more elaborate three-expert AffordanceVLA (95.8%) with under 1% extra parameters. Two probes explain the mechanism: ground-truth affordances fed as an input hurt, and inference-time zeroing shows a lazy bridge acts only as a training-time regularizer while an active bridge becomes load-bearing.
Figures & tables
Fig. 1: Left: our recipe ( Res-Active ). The backbone encodes the camera views and the instruction in a single forward pass; per-layer projections Wlproj hand each layer’s visual tokens v to the corresponding layer of the flow-matching action expert. The affordance head reads uniformly spaced backbone layers through stop-gradient taps, reassembles earlier layers onto finer grids and later layers onto coarser ones, fuses them through a cascade of stages F0,…,FM , and decodes a per-view contact heatmap supervised by Laff . The feature r of an intermediate fusion stage Fk is carried by the bridge Waffproj (written Waff in the text) back into the visual tokens at every conditioning layer. The expert is trained with the action objective Lact in a single pass; at inference it integrates the learned flow over a few passes. Right: the four bridge wirings of our 2×2 ablation, top to bottom Cat-Lazy , Res-Lazy , Cat-Active , and Res-Active (ours). Blue denotes a zero-initialized (lazy) bridge and red a randomly initialized (active) one; a ⊕ joining the adjacent v and r boxes denotes concatenation before the shared projection, whereas a ⊕ on the token stream denotes elementwise addition of the projected r onto the untouched v .
Model
Object
Spatial
Goal
Long
Avg
Base (no head)
98.0
98.6
90.6
85.2
93.1
Head w/o stop-grad †
100.0
92.0
62.0
88.0
85.5
Cat-Lazy
99.4
98.2
94.8
82.6
93.8
Res-Lazy
98.8
96.8
93.4
85.2
93.6
Cat-Active
98.0
99.6
96.0
89.4
95.8
Res-Active
98.8
98.4
96.2
91.4
96.2
TABLE I: LIBERO success rates (%), 50 trials/task. All runs share data, budget, and seed; only the head wiring differs. Bold = best per column. † Measured in a preliminary configuration of the same pipeline.
Fig. 2: Res-Active rollouts with the affordance head’s own predictions overlaid in red. One task per suite, six evenly spaced frames per episode; all four episodes succeed. The predicted contact region locks onto the instruction-designated object among distractors and switches discretely at stage boundaries, e.g. from the stove knob to the moka pot to the burner in the long-horizon task.
Fig. 3: Per-suite success rates of the headless base and the four bridge wirings (50 trials/task). Color encodes topology (blue: concatenation, orange: residual) and saturation encodes initialization (light: lazy, dark: active). The two active wirings dominate on Goal and Long, the suites that stress instruction disambiguation and long-horizon precision, while all five models are near saturation on Object and Spatial.
Fig. 4: The stop-gradient is necessary. The same head helps or hurts depending on this single bit: with the gradient blocked ( Res-Active ) the head lifts the average to 96.2 , while letting the affordance gradient reach the backbone drops it to 85.5 , a 10.7 -point swing. The gradient-exposed variant even falls 7.6 points below the headless base, and the damage is selective: Goal, the suite that leans hardest on instruction grounding, collapses from the 90 s to 62.0 .
Method
Object
Spatial
Goal
Long
Avg
AffordanceVLA (w/o stage II)
91.7
88.5
91.3
73.3
86.2
AffordanceVLA (full)
98.4
98.6
96.2
89.8
95.8
Res-Active (ours)
98.8
98.4
96.2
91.4
96.2
TABLE II: Comparison with AffordanceVLA on LIBERO (%). Their numbers as reported in [ 16 ] .