Memorize, Adapt, Ignore: Diagnosing Robot Learning Mechanisms under Training Data Variation
Authors: Ke Zhang, Danica J. Sutherland, Chao Liu
Organizations: PRIME Robotics Lab, Department of Mechanical Engineering, The University of British Columbia · Department of Computer Science, University of British Columbia · Alberta Machine Intelligence Institute
Training data variation, whether through designing a domain randomization (DR) scheme in simulation or curating demonstrations for imitation learning, is a primary lever for improving the robustness of robotic manipulation policies. Yet its underlying mechanisms remain poorly understood, and practitioners typically select randomization parameters through expensive trial and error. We investigate these mechanisms through a series of case studies, randomizing object size, color, and type as well as scene lighting and linguistic prompts across settings including pick-and-place RL in ManiSkill and fine-tuning of vision-language-action (VLA) models on LIBERO and RoboTwin. We examine both model behavior and internal representations, using the empirical neural tangent kernel (NTK) as our primary diagnostic tool. We show that the NTK distinguishes a shift in the internal learning mechanism from \textit{memorizing} different situations with insufficient variation (e.g.\ learning what to do for a large cube, and what to do for a small cube) to \textit{adapting} to the situation at hand with sufficient variation. An NTK-based signal-to-noise ratio also helps distinguish when policies have learned to \emph{ignore} task-irrelevant factors (e.g.\ treating blue and red cubes identically, instead of learning a blue sub-policy and a red sub-policy). We use these diagnostics to develop practical guidance for designing DR schemes, selecting models, and detecting shortcut learning. We further compare different kinds of representations and validate our findings with real-world hardware experiments using ACT-based imitation learning.
Figures & tables
Figure 1 : Diagnostic metrics tracking representation kernel phase transitions. As the number of randomized colors increases, the model’s learning mechanism transitions from memorizing isolated environments (high effective rank, block-structured eNTK) to adapting across them, and finally to completely ignoring the task-irrelevant nuisance features (collapsing states via PCA, rising SNR, and state CKA alignment) to maximize OOD generalization.
Figure 2 : CNN policies for a ManiSkill pick-cube task.
Table 1 : eNTK SNR tracks OOD success for ACT policies (peg insertion). (a) The no-DR policy has the lowest validation loss but the worst OOD success; both SNRs increase with DR. (b) Joint-probe SNR orders all four training conditions by joint OOD success.
Figure 4 : OpenVLA models evaluated on LIBERO under lighting perturbations.
Figure 5 : Shortcut Learning Detection via Factor Sensitivity Ratio. (a) Training setting illustrating the visual–textual confound in Stage 1 and its resolution in Stage 2. (b) Scene-to-language Factor Sensitivity Ratio (FSR) across training, computed as the representation change induced by scene variation relative to that induced by language variation. The decrease in FSR after introducing the third task indicates reduced relative sensitivity to scene variation and increased sensitivity to language variation.
Figure 6 : Independent variation improves generalization at equal data budget.
Table 2 : Representation kernels have complementary blind spots. (a) The output kernel misses the shortcut because both instructions require the same grasp-frame action. (b) The arms share all layers except their final heads, so the last-hidden kernel cannot see arm-specific structure.
Figure 7 : Lightness domain randomization increases the SNR from roughly 2.5 to 5.5 , indicating features become more separable across the augmented inputs. Samples are annotated with brightness b∈[0.6,1.4] , contrast c∈[0.7,1.3] , and gamma g∈[0.7,1.5] , applied per-sample in that order.
Appendix figures & tables3 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 8 : Real-world evaluation conditions. The Base setting matches training; Bright and Dark are out-of-distribution lighting variations.
Base
Bright (OOD)
Dark (OOD)
Without Lightness DR
6/10
1/10 ( − 83%)
1/10 ( − 83%)
With Lightness DR
10/10
8/10 ( − 20%)
6/10 ( − 40%)
Appendix
Table 3 : Real-world success rates over 10 trials per condition. Parenthesized values give the relative drop from the Base condition. Lightness DR improves OOD success from 1/10 to 8/10 (Bright) and 6/10 (Dark), while also improving in-distribution success.
Similarity space
Spearman ρ(sim(x,x′),Δf(x′))
eNTK K(x,x′)
+0.97±0.06
Activations (best layer)
+0.22±0.28
Activations (MLP head)
+0.09±0.52
Appendix
Table 4 : Predicting training transfer from similarity. Spearman ρ between similarity sim(x,x′) and measured output change Δf(x′) after one SGD step on x , reported as mean ± s.d. across 27 checkpoints ( ≈ 1.7k pairs each). The best activation layer is chosen post hoc.
Visuomotor imitation learning has demonstrated success for manipulation tasks. However, the trained policies remain brittle to visual nuisances', with even minor task-preserving variations such as lighting, distractions or changes in colour result in heavy degradation of the trained policy's performance. While increasing data diversity can improve robustness, it is unclear which additional demonstrations are informative for a particular trained policy. We propose Counterfactual Nuisance Behaviour Cloning (CFNBC), an offline data-selection framework for targeted robustness repair. Starting from a nominal policy trained on clean' demonstrations, CFNBC generates paired clean and nuisance observations that preserve the expert action, then measures \emph{action drift}: the change in the policy's predicted action under a nuisance that should not alter the desired behaviour. This provides a policy-specific sensitivity signal for selecting a compact, response-diverse repair set from a larger candidate pool, without requiring rollout success labels or online policy execution. We show in MuJoCo bimanual cube transfer and SimplerEnv cube stacking that action drift correlates with nuisance-induced failure, and that response-guided repair with only 20--30 selected candidates substantially outperforms matched-budget random selection while approaching the performance of much larger random repair budgets. These results support a data-centric view of robustness repair: the most useful data are not necessarily the most numerous, visually diverse, or obviously difficult, but the examples that cover fragile response modes of the current policy.
Giovanni D'urso, Kaushik Roy, Nicholas Lawrance +1
While Vision-Language-Action (VLA) models offer broad general capabilities, deploying them on specific hardware requires real-world adaptation to bridge the embodiment gap. Since robot demonstrations are costly, this adaptation must often occur under a strict data budget. In this work, we identify a critical diversity trap: the standard heuristic of "maximizing coverage" by collecting diverse, single-shot demonstrations can be self-defeating due to non-vanishing estimation noise. We formalize this phenomenon as a Coverage--Density Trade-off. By decomposing the policy error into estimation (density) and extrapolation (coverage) terms, we characterize an interior optimal allocation of unique conditions for a fixed budget. Guided by this analysis, we propose Anchor-Centric Adaptation (ACA), a two-stage framework that first stabilizes a policy skeleton through repeated demonstrations at core anchors, then selectively expands coverage to high-risk boundaries via teacher-forced error mining and constrained residual updates. Real-robot experiments validate our trade-off framework and demonstrate that ACA significantly improves task reliability and success rates over standard diverse sampling strategies under the same budget.
Yanzhe Chen, Kevin Yuchen Ma, Qi Lv +4
National University of Singapore · Institude for Information Research, A*STAR · Harbin Institute of Technology (Shenzhen).
Vision--language--action (VLA) models acquire broad generalization through large-scale pretraining, yet adapting them to a new task and robot embodiment still requires post-training on newly collected data. Unlike pretraining, post-training targets task- and embodiment-specific adaptation, making it particularly sensitive to data quality. In practice, collected robot datasets often contain heterogeneous errors, including execution mistakes, sensor drift, and timestamp misalignment, which can impair post-training and policy performance. Manual inspection is costly, while existing data-cleaning methods are typically tailored to particular corruption types. To address these challenges, we introduce \textsc{RoboDrop}, a data-curation framework that audits supervision using local gradient compatibility measured along the training trajectory as a proxy for its effect on post-training performance. During a one-epoch warm-up run, RoboDrop scores each candidate sample online by comparing its gradient with those of task-semantic and visually matched validation samples. The resulting sample scores are aggregated at the episode level, and a simple automatic post-processing rule converts them into filtering decisions. We evaluate RoboDrop on controlled observation--action corruptions, naturally suboptimal demonstrations in simulation, and real-robot datasets containing non-expert collection errors. Across these settings, RoboDrop more accurately distinguishes unreliable demonstrations than prior methods, while post-training on the curated data consistently yields stronger downstream policies, with average real-robot rollout success rising from 35.0% to 67.5%. These results establish training-trajectory-aware, context-conditioned supervision auditing as an effective approach to robust VLA post-training.