While pretrained robotic policies exhibit impressive capabilities in controlled environments, unobserved physical properties and dynamics require these policies to rapidly adapt during deployment. Existing test-time adaptation methods typically rely on sparse scalar rewards, failing to exploit the rich geometric and dynamic feedback from the environment during physical interaction. To address this challenge, we propose SCOUT, a dynamics-aware meta-learning framework that enables manipulation policies to rapidly adapt by continuously revising their internal beliefs about environment dynamics. Our approach couples an action-prediction policy with a forward dynamics model via a shared belief latent space. During meta-training, an inner loop updates this shared belief latent by minimizing the dynamics prediction error against the observed action outcome, while the outer loop optimizes the network for action selection. At deployment, this structure allows the agent to infer and adapt to unknown physical dynamics on the fly. By updating its latent belief based on action-outcome mismatches, the policy automatically adapts without risking catastrophic forgetting. We demonstrate that SCOUT significantly accelerates online adaptation across simulated manipulation benchmarks and achieves robust sim-to-real transfer in the real world. Project webiste can be found here: https://liy1shu.github.io/SCOUT/
Figures & tables
Figure 2: Overview of our test-time adaptation framework. In the inner loop, the initial belief q0 is adapted to qk using past interactions by minimizing the latent dynamics prediction error ( Linner ) while keeping all network weights frozen. In the outer loop, the adapted belief qk modulates the policy head via FiLM to predict the action distribution. The entire framework, including q0 , is meta-trained end-to-end by backpropagating the policy loss ( Louter ) through the inner loop updates.
Method
Slide Brick
Push Bar
Pick Bar
Open Box
Turn Faucet
All (Normalized)
CQL [ 42 ]
9.94 ± 0.18
10.96 ± 0.20
10.22 ± 0.19
4.32 ± 0.18
6.63 ± 0.21
3.17 ± 0.04
BC
10.24 ± 0.33
11.26 ± 0.30
9.79 ± 0.32
2.59 ± 0.19
4.54 ± 0.27
2.72 ± 0.05
ADVC [ 43 ]
8.36 ± 0.27
7.83 ± 0.27
5.41 ± 0.22
2.82 ± 0.16
2.67 ± 0.16
1.91 ± 0.04
ISE [ 40 ]
7.69 ± 0.26
5.26 ± 0.20
4.84 ± 0.18
2.25 ± 0.10
2.39 ± 0.14
1.61 ± 0.03
PEARL [ 13 ]
5.03 ± 0.23
4.51 ± 0.20
4.89 ± 0.25
2.24 ± 0.12
1.54 ± 0.03
1.31 ± 0.03
Raw GMM( B.3 )
6.39 ± 0.28
7.95 ± 0.25
4.96 ± 0.20
2.41 ± 0.10
2.22 ± 0.07
1.67 ± 0.03
Table 1: ISE benchmark results . The average number of attempts (along with standard errors) needed for success across 400 rollouts. Lower numbers denote faster adaptation. Normalized score refers to an aggregated efficiency score that compares all methods against the performance of our method. The best performing method is bolded. SCOUT achieves top efficiency across all tasks.
Generator Only
Generator + Verifier
SCOUT
PNDiT [ 31 ]
Raw GMM
History-cond. GMM
HAVE [ 31 ] (5 samples)
HAVE [ 31 ] (20 samples)
SCOUT (1 step)
SCOUT (5 steps)
Success Rate
0.68
0.61
0.68
0.82
0.84
0.83
0.97
Avg Steps
15.0
15.8
18.7
13.4
12.9
13.6
12.8
Time (s)
0.463
0.002
0.004
0.757
1.876
0.012
0.052
Table 2: Ambiguous door results . SCOUT outperforms baselines without the high time cost required by generating multiple samples. (# steps) for SCOUT indicates the number of adaptation steps used.
Figure 3: Real world uneven bar pickup and push results . Left: The number of attempts to success with and without test-time training, reported with 95% confidence interval. Right: A visualization of the policy prediction weights across the x-axis of the uneven bar before and after adaptation. We can see that SCOUT efficiently adapts to different Center of Mass positions.
Figure 4: Belief Latent Space Evolution. PCA projection of the belief latent qk across three stages of adaptation. Points are colored by their ground-truth motion mode, while background regions represent kNN-estimated decision boundaries. As test-time training progresses, the latents successfully migrate from the shared prior q0 into distinct, well-separated kinematic clusters.
Figure 5: Evolution of predicted action distributions across rollout steps . (Top) Pick Bar: Despite high initial variance, the predicted grasp distribution quickly converges toward the optimal center of mass. (Bottom) Open Door: Probability mass across discrete action modes successfully concentrates onto the correct target action (i.e., Pull Right) as online adaptation progresses.
Trained
Held-out
Raw GMM
Task
R1/R3/R5
Avg
R2/R4
Avg
(full range)
Pick Bar
2.8/2.2/1.7
2.25
4.0/5.1
4.56
4.96
Push Bar
3.7/5.7/3.9
4.48
5.9/6.8
6.34
7.95
Slide Brick
4.6/3.0/2.1
3.14
3.0/3.7
3.36
6.39
Table 3: Held-out dynamics ranges. Average attempts to success when training on parameter segments R1/R3/R5 and evaluating on all five. The last column is the raw GMM trained on the full range (Table 1 ).
Appendix figures & tables7 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 6: PCA projection of OutcomeEncoder latents z on 120 validation outcomes , with (left) and without (right) the reconstruction loss, drawn on shared axes. Because the encoder’s projection head ends in a LayerNorm , every z has a near-constant norm ( ∥z∥≈13 – 20 ), so collapse manifests as a common direction rather than a vanishing magnitude. With Lrecon the latents span a wide arc (per-axis spread std 7.9 ), indicating that outcome-specific information is retained in z . Without it the latents collapse to a single point at the origin (spread std 0.01 ), a 644× reduction in mean pairwise distance relative to the reconstruction model.
SCOUT (w/o Recon. loss)
SCOUT (Full)
1 step
5 steps
1 step
5 steps
Success Rate
0.71
0.73
0.83
0.97
Avg Steps
14.2
14.7
13.6
12.8
Appendix
Table 4: Reconstruction loss ablation on the ambiguous doors . Success Rate ( ↑ ) and Avg. Steps ( ↓ ) for SCOUT with and without the reconstruction loss, each evaluated with 1 and 5 adaptation steps. Bold denotes the best result per metric.
Median
Mean
p
Task
SCOUT
Baseline
SCOUT
Baseline
U
raw
Holm
Psup
Pick Bar
2
3
2.17
3.27
311.5
0.035
0.035
0.65
Push Bar
3
4
2.90
5.60
250.0
0.003
0.005
0.72
Appendix
Table 5: Real-world significance tests, pooled over CoM positions. Two-sided Mann–Whitney U with tie and continuity correction, n=30 rollouts per method. Psup is the probability that a random SCOUT rollout needs fewer attempts than a random baseline rollout.
Estimated flow
Ground-truth flow
1 step
5 steps
1 step
5 steps
Success Rate
0.83
0.96
0.83
0.97
Avg Steps
15.86
14.94
13.6
12.8
Appendix
Table 6: Robustness to estimated flow on the ambiguous doors . Success Rate ( ↑ ) and Avg. Steps ( ↓ ) for SCOUT when the action outcome is the simulator’s ground-truth flow (as in Table 2 ) or flow estimated at test time.
Belief used on door B
NLL ( ↓ )
Top-100 acc. ( ↑ )
No adaptation ( q=q0 )
6.41
0.26
Adapted on B itself
3.62
0.89
Adapted on a different door A (same mode)
3.59
0.91
Appendix
Table 7: Cross-trajectory belief transfer on the ambiguous doors . The policy is evaluated on query door B using a belief adapted on B itself, on a different door A with the same articulation mode, or not adapted at all.
Figure 7: Decoded center of mass during real-world adaptation. A linear probe decodes the CoM position from the belief latent at the first, middle, and final attempt of each real-world rollout; color denotes the true CoM configuration. Decoded positions start near a shared prior and separate toward their true values as adaptation proceeds.
Figure 8: Real World Setup. The Franka Emika Panda robot is equipped with a Franka Gripper and observes the workspace through a ZED camera.
School of Computer Science and Engineering, Tongji University, Shanghai, China · School of Vehicle and Mobility, Tsinghua University, Beijing, China · Simple AI, Beijing, China +1