cs.AIOct 7, 2026

Self-Supervised Keyframe Discovery for Horizon-Invariant Behavior Cloning

Authors: Prabin Kumar Rath, Omkar Patil, Nakul Gopalan

Organizations: Arizona State University

Abstract

Behavior cloning (BC) in non-Markovian environments is a challenging problem because policies have to reason over contextual information over long horizons. Existing policy architectures rely on recurrent or attention-based mechanisms to capture long-term dependencies. However, recurrent models suffer from hidden-state collapse and gradient instability under backpropagation through time, while attention-based models are fundamentally limited by context length. To address these issues, we propose Keyframe Mnemonics, a novel self-supervised method that discovers\textit{discovers} a set of information-critical observations (mnemonics\textit{mnemonics}) by learning an objective from randomly sampled past observations and using it as a reward for keyframe selection. We then train a BC policy that conditions on the discovered keyframes to model the action distribution. Under certain task-structure assumptions, our formulation provides context retention guarantees over an infinite horizon, while maintaining a small set of decision-relevant keyframes in the policy's working memory. We evaluate our method on synthetic memory domains, where mnemonic-conditioned BC policies achieve 100100% success rates (SR) and generalize to horizons orders of magnitude beyond training without performance degradation. Additionally, we evaluate on memory-intensive robot manipulation benchmark, achieving a 13.913.9% average absolute SR improvement over the strongest baseline across 2323 tasks and retaining 8080% SR at 20×20\times longer horizons on a real robot. Code and videos are available at https://keyframe-mnemonics.github.io.

Figures & tables

Appendix figures & tables27 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Training-free Behavior Cloning

    Sep 24, 2026Maximilian Adang, Timothy Chen, Lars Osterberg +2Robot Policy LearningBehavior Cloning

  2. Difference-Aware Retrieval Policies for Imitation Learning

    Jun 8, 2026Quinn Pfeifer, Ethan Pronovost, Paarth Shah +3Imitation LearningInteractive Imitation Learning

  3. Scalable Behavior Cloning with Open Data, Training, and Evaluation

    Jun 25, 2026Arthur Allshire, Himanshu Gaurav Singh, Ritvik Singh +15Robot Policy LearningDiffusion Transformer