Credit-Guided Policy Improvement for Test-time Adaptive Vision-Language Navigation
Authors: Yang Li, Sijia Zhang, Yihan Li, Aming WU, Zihao Zhang, Ziju Han, Yahong Han
Organizations: School of Artificial Intelligence, Tianjin University, China · The University of Hong Kong, Hong Kong SAR, China · School of Computer Science and Information Engineering, Hefei University of Technology, China
Test-time adaptation for vision-language navigation (TTA-VLN) enables pretrained policies to adapt online to unseen environments using only test-time observations and interaction history. However, distribution shifts can distort local action preferences and lead to off-course decisions. Existing methods rely on predictive uncertainty, trajectory-level feedback, or accumulated adaptation experience to correct such deviations. These signals, however, do not directly reveal whether an executed action supports instruction-guided progress toward the goal. Moreover, a plausible corrective signal does not guarantee a reliable policy update. The key challenge is thus twofold: identifying interactions that support goal-directed improvement and determining whether the resulting updates are worth retaining. We observe that each executed action induces an immediate observation transition, providing evidence of its local consequences. Based on this insight, we propose Credit-Guided Policy Improvement (CGPI), which recovers signed, reference-relative decision credit from action-induced observation transitions without external outcome feedback. With the pretrained navigation policy frozen, CGPI uses this credit to propose lightweight adaptation updates and verifies them against prior credit-supported interactions. Updates are retained only when supported and rolled back otherwise. CGPI achieves consistent gains across the evaluated VLN benchmarks and navigation backbones, while qualitative robot trials further illustrate the feasibility of zero-shot sim-to-real transfer.
Figures & tables
Figure 1: Motivation for CGPI. TTA enables online policy change, but change does not necessarily imply improvement. Decision credit first identifies whether the current interaction provides local evidence of goal-directed improvement; the resulting candidate update is then verified against past credit-supported interactions before being retained or rolled back.
Figure 2: Overview of CGPI. The method constructs local process-evidence hypotheses and grounds the executed decision using its actual resulting observation. Reference-relative comparison recovers signed decision credit, which proposes a candidate modification to an action reranker. The candidate is then verified against past credit-supported interactions and retained only when sufficiently supported and satisfying the proposal-state reference constraint; otherwise, it is rolled back.
Methods+Model
REVERIE Val Seen
REVERIE Val Unseen
REVERIE Test Unseen
OSR ↑
SR ↑
SPL ↑
RGSPL ↑
OSR ↑
SR ↑
SPL ↑
RGSPL ↑
OSR ↑
SR ↑
SPL ↑
RGSPL ↑
HAMT [ 4 ]
47.65
43.29
40.19
25.18
36.84
32.95
30.20
17.28
33.41
30.40
26.67
13.08
+ Tent [ 18 ]
46.03
43.43
40.78
25.81
32.60
30.56
28.23
14.48
25.06
23.73
21.78
10.82
+ SAR [ 16 ]
47.12
43.85
39.97
25.44
32.86
31.12
29.15
15.23
27.94
25.86
23.08
11.58
+ ViDA [ 15 ]
47.67
43.55
41.23
25.27
32.74
30.97
28.82
14.97
27.03
24.81
22.45
11.23
+ FSTTA [ 7 ]
48.21
42.87
39.56
24.58
36.78
32.89
30.51
17.20
33.39
30.39
26.65
13.61
Table 1: Experimental results for different TTA strategies on the REVERIE dataset.
Table 4
Figure 3: Real-world qualitative results of CGPI on a Unitree Go2 robot. We show third-person views and egocentric observations for three executed trajectories.
Figure 4: Qualitative comparison of IDEA and CGPI in unseen environments.
Variant
Grounding
Reference
SR ↑
SPL ↑
CPC ↑
WUR ↓
Absolute Hypothesis
–
–
54.32
36.71
0.297
0.312
Grounded Evidence
✓
–
56.14
38.06
0.414
0.244
Relative Hypothesis
–
✓
56.82
38.67
0.463
0.221
CGPI
✓
✓
58.76
40.15
0.572
0.154
Table 4: Ablation of CGPI’s two core improvement-identification operations on REVERIE Val Unseen with DUET.
Movement Credit
STOP Credit
SR ↑
SPL ↑
–
–
30.40
26.67
✓
–
33.54
29.61
–
✓
31.46
27.55
✓
✓
34.73
30.78
Table 5: Contribution of movement-action credit and STOP-specific credit in CGPI on REVERIE.
Variant
Verify
Cand. KL
SR ↑
SPL ↑
PD ↓
AIR ↑
RR(%)
CGPI w/o Verification
–
–
57.41
38.96
0.071
0.548
100.0
CGPI w/o Candidate KL
✓
–
58.29
39.72
0.057
0.704
67.8
Full CGPI
✓
✓
58.76
40.15
0.043
0.741
61.3
Table 6: Ablation of improvement verification and retention on REVERIE Val Unseen with DUET.
Method
AIR ↑
AIG (m) ↑
RP ↑
Credit-Guided Direct Commit
0.548
0.19
–
CGPI
0.741
0.63
0.716
Table 7: Policy-improvement analysis on REVERIE. Left: post-hoc candidate-update quality (DUET, Val Unseen). Right: within-episode versus continual adaptation (HAMT, Test Unseen).