Credit-Guided Policy Improvement for Test-time Adaptive Vision-Language Navigation
Authors: Yang Li, Sijia Zhang, Yihan Li, Aming WU, Zihao Zhang, Ziju Han, Yahong Han
Organizations: School of Artificial Intelligence, Tianjin University, China · The University of Hong Kong, Hong Kong SAR, China · School of Computer Science and Information Engineering, Hefei University of Technology, China
Test-time adaptation for vision-language navigation (TTA-VLN) enables pretrained policies to adapt online to unseen environments using only test-time observations and interaction history. However, distribution shifts can distort local action preferences and lead to off-course decisions. Existing methods rely on predictive uncertainty, trajectory-level feedback, or accumulated adaptation experience to correct such deviations. These signals, however, do not directly reveal whether an executed action supports instruction-guided progress toward the goal. Moreover, a plausible corrective signal does not guarantee a reliable policy update. The key challenge is thus twofold: identifying interactions that support goal-directed improvement and determining whether the resulting updates are worth retaining. We observe that each executed action induces an immediate observation transition, providing evidence of its local consequences. Based on this insight, we propose Credit-Guided Policy Improvement (CGPI), which recovers signed, reference-relative decision credit from action-induced observation transitions without external outcome feedback. With the pretrained navigation policy frozen, CGPI uses this credit to propose lightweight adaptation updates and verifies them against prior credit-supported interactions. Updates are retained only when supported and rolled back otherwise. CGPI achieves consistent gains across the evaluated VLN benchmarks and navigation backbones, while qualitative robot trials further illustrate the feasibility of zero-shot sim-to-real transfer.
Figures & tables
Figure 1: Motivation for CGPI. TTA enables online policy change, but change does not necessarily imply improvement. Decision credit first identifies whether the current interaction provides local evidence of goal-directed improvement; the resulting candidate update is then verified against past credit-supported interactions before being retained or rolled back.
Figure 2: Overview of CGPI. The method constructs local process-evidence hypotheses and grounds the executed decision using its actual resulting observation. Reference-relative comparison recovers signed decision credit, which proposes a candidate modification to an action reranker. The candidate is then verified against past credit-supported interactions and retained only when sufficiently supported and satisfying the proposal-state reference constraint; otherwise, it is rolled back.
Methods+Model
REVERIE Val Seen
REVERIE Val Unseen
REVERIE Test Unseen
OSR ↑
SR ↑
SPL ↑
RGSPL ↑
OSR ↑
SR ↑
SPL ↑
RGSPL ↑
OSR ↑
SR ↑
SPL ↑
RGSPL ↑
HAMT [ 4 ]
47.65
43.29
40.19
25.18
36.84
32.95
30.20
17.28
33.41
30.40
26.67
13.08
+ Tent [ 18 ]
46.03
43.43
40.78
25.81
32.60
30.56
28.23
14.48
25.06
23.73
21.78
10.82
+ SAR [ 16 ]
47.12
43.85
39.97
25.44
32.86
31.12
29.15
15.23
27.94
25.86
23.08
11.58
+ ViDA [ 15 ]
47.67
43.55
41.23
25.27
32.74
30.97
28.82
14.97
27.03
24.81
22.45
11.23
+ FSTTA [ 7 ]
48.21
42.87
39.56
24.58
36.78
32.89
30.51
17.20
33.39
30.39
26.65
13.61
Table 1: Experimental results for different TTA strategies on the REVERIE dataset.
Table 4
Figure 3: Real-world qualitative results of CGPI on a Unitree Go2 robot. We show third-person views and egocentric observations for three executed trajectories.
Figure 4: Qualitative comparison of IDEA and CGPI in unseen environments.
Variant
Grounding
Reference
SR ↑
SPL ↑
CPC ↑
WUR ↓
Absolute Hypothesis
–
–
54.32
36.71
0.297
0.312
Grounded Evidence
✓
–
56.14
38.06
0.414
0.244
Relative Hypothesis
–
✓
56.82
38.67
0.463
0.221
CGPI
✓
✓
58.76
40.15
0.572
0.154
Table 4: Ablation of CGPI’s two core improvement-identification operations on REVERIE Val Unseen with DUET.
Movement Credit
STOP Credit
SR ↑
SPL ↑
–
–
30.40
26.67
✓
–
33.54
29.61
–
✓
31.46
27.55
✓
✓
34.73
30.78
Table 5: Contribution of movement-action credit and STOP-specific credit in CGPI on REVERIE.
Variant
Verify
Cand. KL
SR ↑
SPL ↑
PD ↓
AIR ↑
RR(%)
CGPI w/o Verification
–
–
57.41
38.96
0.071
0.548
100.0
CGPI w/o Candidate KL
✓
–
58.29
39.72
0.057
0.704
67.8
Full CGPI
✓
✓
58.76
40.15
0.043
0.741
61.3
Table 6: Ablation of improvement verification and retention on REVERIE Val Unseen with DUET.
Method
AIR ↑
AIG (m) ↑
RP ↑
Credit-Guided Direct Commit
0.548
0.19
–
CGPI
0.741
0.63
0.716
Table 7: Policy-improvement analysis on REVERIE. Left: post-hoc candidate-update quality (DUET, Val Unseen). Right: within-episode versus continual adaptation (HAMT, Test Unseen).
Test-time training (TTT) offers a lightweight way to adapt vision--language--action (VLA) policies from unlabeled deployment streams, but it remains difficult to use reliably in closed-loop manipulation. A shared adaptation space can mix incompatible task corrections, while an online update can alter subsequent actions before its consequences are known. We introduce a reliable TTT framework for VLA policies (VANE). VANE conditions prompt adaptation on the current vision--language context and learns from the future visual consequences of executed actions. Candidate updates are isolated from the live policy, evaluated on subsequent observations, and committed only when supported by future evidence, making adaptation selective and reversible. On SimplerEnv WidowX, VANE improves average success by 3.2 percentage points over the corresponding TTT baseline. Results on Google Robot further show that deployment-time gains remain task- and embodiment-dependent. Together, these results demonstrate a constrained, evidence-based approach to adapting VLA policies during interaction.
Hongjin Ji, Guoyang Xia, Luoyang Sun +2
The Chinese University of Hong Kong, Shenzhen · Li Auto Inc. · Beijing University of Posts and Telecommunications +2
Vision-Language Navigation (VLN) requires embodied agents to generate actions based on instructions and observations. General-purpose multimodal agents offer a promising basis for this task, but selecting plausible local actions does not ensure that execution remains consistent with the intended route, particularly in long-horizon tasks. Moreover, the accumulated interaction history increases the input required for subsequent decisions, resulting in a significant inference overhead. To this end, we introduce \method, an Agentic VLN framework that includes a Goal Agent that sets adaptive goals for local actions, a Verify Agent that dynamically verifies whether a goal has been completed, a Memory Agent for multimodal context compression, and a Visuomotor Agent to execute adaptive goals. Specifically, the Goal Agent formulates adaptive goals based on the instruction, current observation, and execution history. Then the Visuomotor Agent executes navigation actions to achieve each goal, while the Verify Agent uses a goal-specific verification question to dynamically assess whether the observed outcomes satisfy the intended completion condition. Verified goal completion then marks a boundary for the Memory Agent to compress the corresponding multimodal interaction history while preserving information needed for subsequent navigation. We evaluate navigation on R2R-CE and RxR-CE, examine framework variants across three model backbones, and study context evolution during execution. For Real-World evaluation, \method achieves 83.3% success and 1.51,m navigation error across eight challenging routes evaluated three times each.
Haoxiang Shi, Zaijing Li, Muhe Ding +3
Harbin Institute of Technology (Shenzhen) · Pengcheng Laboratory
Existing vision-language navigation methods often couple a VLM with waypoint decoders to produce multi-step action plans, but they typically lack an explicit closed-loop mechanism for tracking semantic progress, diagnosing execution failures, and recovering from error accumulation in long-horizon navigation. To address this gap, we propose ReflectVLN, an agentic VLN framework that organizes decision-making through bidirectionally interactive intention and execution agents. The intention agent performs subtask decomposition and reflection, generating executable subtask descriptions as corrective plans. Conditioned on these descriptions, the execution agent grounds them into short-horizon actions under current observations while monitoring sub-goal progress and detecting off-track behavior. Crucially, ReflectVLN enables closed-loop bidirectional communication: the execution agent emits progress and deviation signals to trigger reflection and subtask updates on demand, and the intention agent returns structured guidance that reconditions subsequent actions for recovery. To encourage temporally coherent decisions with interpretable intermediate rationales, we introduce Action Chain-of-Thought (Action-CoT), a path-conditioned dual-query training scheme for action generation. Experiments on standard VLN benchmarks show that ReflectVLN improves success rates and path efficiency under a constrained data budget, with favorable training cost and fewer high-level intention calls at inference time, while providing interpretable intermediate decisions for analysis and collaboration. Code is available at: https://github.com/AIprogrammer/ReflectVLN