Wiring Matters: Injection Topology and Initialization of Affordance Heads in Vision-Language-Action Policies
Authors: Zijian An, Linhan Wang, Jiayan Wang, Shijie Geng, Ran Yang, Yiming Feng, Lifeng Zhou
Organizations: Drexel University, Philadelphia, PA 19104, USA · Virginia Tech, Virginia Seafood Agricultural Research and Extension Center, Hampton, VA 23669, USA · TODO: Jiayan Wang’s affiliation and email · Amazon Store Foundation AI (SFAI), New York, NY 10018, USA
Dense affordance supervision is an appealing auxiliary signal for vision-language-action (VLA) policies, yet naively co-training an affordance head can severely damage instruction following. We present a controlled study of how to wire such a head into a modern VLA on the LIBERO benchmark. Our recipe reads the backbone through a stop-gradient and re-injects an intermediate head feature into the action expert via a learned bridge. The stop-gradient is a precondition: letting affordance gradients reach the backbone drops the policy below the headless base (85.5% vs. 93.1%). With the backbone protected, a same-budget 2*2 ablation over injection topology (concatenation vs. residual) and bridge initialization (zero vs. random) shows initialization is the dominant lever. The best wiring, an actively initialized residual bridge, reaches 96.2%, matching the far more elaborate three-expert AffordanceVLA (95.8%) with under 1% extra parameters. Two probes explain the mechanism: ground-truth affordances fed as an input hurt, and inference-time zeroing shows a lazy bridge acts only as a training-time regularizer while an active bridge becomes load-bearing.
Figures & tables
Fig. 1: Left: our recipe ( Res-Active ). The backbone encodes the camera views and the instruction in a single forward pass; per-layer projections Wlproj hand each layer’s visual tokens v to the corresponding layer of the flow-matching action expert. The affordance head reads uniformly spaced backbone layers through stop-gradient taps, reassembles earlier layers onto finer grids and later layers onto coarser ones, fuses them through a cascade of stages F0,…,FM , and decodes a per-view contact heatmap supervised by Laff . The feature r of an intermediate fusion stage Fk is carried by the bridge Waffproj (written Waff in the text) back into the visual tokens at every conditioning layer. The expert is trained with the action objective Lact in a single pass; at inference it integrates the learned flow over a few passes. Right: the four bridge wirings of our 2×2 ablation, top to bottom Cat-Lazy , Res-Lazy , Cat-Active , and Res-Active (ours). Blue denotes a zero-initialized (lazy) bridge and red a randomly initialized (active) one; a ⊕ joining the adjacent v and r boxes denotes concatenation before the shared projection, whereas a ⊕ on the token stream denotes elementwise addition of the projected r onto the untouched v .
Model
Object
Spatial
Goal
Long
Avg
Base (no head)
98.0
98.6
90.6
85.2
93.1
Head w/o stop-grad †
100.0
92.0
62.0
88.0
85.5
Cat-Lazy
99.4
98.2
94.8
82.6
93.8
Res-Lazy
98.8
96.8
93.4
85.2
93.6
Cat-Active
98.0
99.6
96.0
89.4
95.8
Res-Active
98.8
98.4
96.2
91.4
96.2
TABLE I: LIBERO success rates (%), 50 trials/task. All runs share data, budget, and seed; only the head wiring differs. Bold = best per column. † Measured in a preliminary configuration of the same pipeline.
Fig. 2: Res-Active rollouts with the affordance head’s own predictions overlaid in red. One task per suite, six evenly spaced frames per episode; all four episodes succeed. The predicted contact region locks onto the instruction-designated object among distractors and switches discretely at stage boundaries, e.g. from the stove knob to the moka pot to the burner in the long-horizon task.
Fig. 3: Per-suite success rates of the headless base and the four bridge wirings (50 trials/task). Color encodes topology (blue: concatenation, orange: residual) and saturation encodes initialization (light: lazy, dark: active). The two active wirings dominate on Goal and Long, the suites that stress instruction disambiguation and long-horizon precision, while all five models are near saturation on Object and Spatial.
Fig. 4: The stop-gradient is necessary. The same head helps or hurts depending on this single bit: with the gradient blocked ( Res-Active ) the head lifts the average to 96.2 , while letting the affordance gradient reach the backbone drops it to 85.5 , a 10.7 -point swing. The gradient-exposed variant even falls 7.6 points below the headless base, and the damage is selective: Goal, the suite that leans hardest on instruction grounding, collapses from the 90 s to 62.0 .
Method
Object
Spatial
Goal
Long
Avg
AffordanceVLA (w/o stage II)
91.7
88.5
91.3
73.3
86.2
AffordanceVLA (full)
98.4
98.6
96.2
89.8
95.8
Res-Active (ours)
98.8
98.4
96.2
91.4
96.2
TABLE II: Comparison with AffordanceVLA on LIBERO (%). Their numbers as reported in [ 16 ] .
Recent advances in Vision-Language-Action (VLA) models have shown strong potential for general-purpose robotic manipulation. However, the visual representations of most VLA models are often dominated by global object appearance and struggle to focus on task-relevant functional interaction regions, which limits their robustness in unstructured environments. Existing affordance-based methods typically rely on explicit mask injection or external perception modules, requiring additional annotations while introducing cascading perception errors and inference overhead. To address these limitations, we propose AffordVLA, an affordance-enhanced VLA framework that internalizes manipulation-centric affordance perception into VLA visual representations through implicit representation alignment. Specifically, we construct a zero-shot affordance teacher to extract task-conditioned affordance visual representations from RGB observations and language instructions. AffordVLA aligns the intermediate visual representations of the VLA with the affordance visual representations extracted by the teacher, thereby implicitly injecting manipulation-centric affordance perception into VLA visual representations and improving action accuracy. Extensive simulation and real-world experiments demonstrate that AffordVLA and its affordance teacher achieve state-of-the-art performance and outperform strong baselines. Ablation analyses show that AffordVLA effectively reshapes VLA visual representations while preserving inference efficiency, leading to improved manipulation success rates and training efficiency.
Weijie Kong, Zhian Su, Wei Yu +1
Grasp Lab, School of Mechanical Engineering of Zhejiang University, Hangzhou 310027, China · Torch Kernel Co., Ltd., Hangzhou, 310000, P. R. China
Current Vision-Language-Action (VLA) models typically treat the deepest representation of a vision-language backbone as universally optimal for action prediction. However, robotic manipulation is composed of many frequent closed-loop spatial adjustments, for which excessive abstraction may waste computation and weaken low-level geometric cues essential for precise control. Existing early-exit strategies attempt to reduce computation by stopping at predefined layers or applying heuristic rules such as action consistency, but they do not directly answer when a representation is actually sufficient for action. In this paper, we present LoopVLA, a recurrent VLA architecture that jointly learns representation refinement, action prediction, and sufficiency estimation. LoopVLA iteratively applies a shared Transformer block to refine multimodal tokens, and at each iteration produces both a candidate action and a sufficiency score that estimates whether further refinement is necessary. By sharing parameters across iterations, LoopVLA decouples refinement from absolute layer indices and grounds sufficiency estimation in the evolving representation itself. Since sufficiency has no direct supervision, we introduce a self-supervised distribution alignment objective, where intermediate confidence scores are trained to match the relative action quality across refinement steps, thereby linking sufficiency learning to policy optimization signals. Experiments on LIBERO, LIBERO-Plus, and VLA-Arena show that LoopVLA pushes the efficiency-performance frontier of VLA policies, reducing parameters by 45% and improving inference throughput by up to 1.7 times while matching or outperforming strong baselines in task success.
Boyang Shen, Kaixiang Yang, Hao Wang +4
1Huazhong University of Science and Technology · 2Wuhan United Imaging Surgical Co.,Ltd. (UIS)
Vision-language-action (VLA) policies typically feed dense visual patch tokens into a language-action backbone, preserving scene context but offering no explicit mechanism to regulate how strongly different visual tokens influence policy computation. We introduce DIVA, a Dual-Space Intent-Aware Visual Attenuation module with an anchor-then-attenuate design. DIVA combines high-level task intent with low-level visual evidence to estimate patch-wise relevance anchors, then applies them in two complementary spaces: it reweights projected visual tokens before backbone entry and persistently attenuates low-relevance visual states within the backbone. DIVA preserves the full visual token sequence and requires no external grounding supervision. On LIBERO, DIVA improves OpenVLA-OFT from 96.6% to 98.0% average success and raises its zero-shot LIBERO-Plus score from 69.6 to 72.6. Real-world experiments further show consistent gains under task-irrelevant visual perturbations, supporting the robustness of intent-aware visual attenuation beyond simulation.