Neural behavior cloning compresses demonstrations into large models, making individual actions difficult to trace and policy updates costly. Retrieval policies retain access to demonstrations but struggle with mismatch between recorded and live behavior. We introduce Behavior Predictive Control (BPC), which synthesizes policies without end-to-end policy training by combining an action-aware retrieval metric, a Hankel-based action-continuation prior, and a closed-form one-step residual correction. Inspired by behavioral systems theory, BPC predicts future actions by blending stored observation-action data that best reconstructs the recent runtime observation--action history. Across simulated benchmarks and real-robot deployments, BPC is competitive with learned policies such as π0.5 (surpassing it in some cases), while reducing policy fitting from hours to seconds on consumer GPUs and supporting closed-loop control upwards of 75 Hz on a Jetson Orin Nano. The retrieved demonstration windows and their coefficients also provide an intrinsic estimate of task progress. Retaining demonstrations within the deployed policy makes its predictions traceable to supporting trajectories and enables behavior revision through the demonstration bank.
Figures & tables
Policy family
Fit time
Inference
Memory
Pretraining?
Performance
Interpretable?
Learned [ 1 , 13 , 5 ]
Hours
Slow; real-time may need action chunking
High; may need server GPUs
Frequently
Strong
Opaque
Retrieval [ 7 , 8 ]
Minutes
Fast
Low
Pretrained encoder required
Lower
High
BPC (ours)
Seconds
Fast; 75 Hz on Jetson Orin Nano
Low
Optional; raw pixels may suffice
Strong
High
TABLE I : Policy archetype tradeoffs.
Fig. 2 : Online BPC pipeline. (1) A learned metric retrieves K demonstration windows from the live observation–action history z . (2) Regularized reconstruction fits coefficients gi to their past histories and applies them to future action blocks to form the action prior. (3) Random Fourier features of retrieved observations, actions, and observation differences from the query are uniformly pooled. A linear head predicts the correction δu(t) . The action prior is added to the correction to form the outgoing control action.
TABLE II : Tabletop manipulation hardware: qualitative task sequences and successes over number of trials.
TABLE III : Success rates (%) on simulation tabletop manipulation tasks over 300 rollouts per task.
Task
Success rate
Collision rate (c/m)
Demo-set distance (m)
Demos / demo duration (s)
Loop
10/10
0.0000
0.095
3 / 25.2
Figure-eight
8/10
0.0108
0.109
4 / 33.5
TABLE IV : Drone hardware results for Loop and Figure-eight . Demonstration-set distance is the mean 3D distance to the nearest demonstration trajectory. Demo duration is the mean duration per demonstration.
Fig. 3 : Representative hardware trajectory snapshots for Loop (left) and Figure-eight (right). Demonstrations, flown paths, and predicted references are shown alongside the preceding two second history. Colored demonstration segments indicate normalized coefficient magnitudes, ∣gi∣/∑j∣gj∣ .
TABLE V : Success rates (%) on the dexterous manipulation suites with 120 rollouts per task: DexArt [ 37 ] and Adroit [ 38 ] .
Method
Square
Stack
Coffee
Hammer
Mug
Nut
Stack3
Thread
Mean
NHC
76
88
91
98
63
4
23
56
62
NNC
73
10
91
96
31
9
0
50
45
NLR
71
91
94
99
68
9
27
64
65
BPC (Ours)
73
96
96
99
78
13
50
77
73
TABLE VI : Ablations on tabletop manipulation tasks: success rate (%) over 300 trials per task.
Fig. 5 : Computation scaling across different policies (BPC and π0.5 ) across the number of windows in the bank. Left: fitting time and per-call execution latency (log scale); right: peak GPU memory. Solid lines and filled shapes show fitting, while dotted lines and open shapes show test-time values. Squares mark π0.5 on an RTX 4090, and circles represent BPC values on an RTX 4090.
Setting
MimicGen
Adroit
DexArt
Demonstrations per task
1,000
10
100
History Tini
10 (Square: 5)
10
10
Continuation Tf
10
10
10
Actions executed per update
1
1
1
Warmup
10 (Square: 5)
10
10
Maximum episode length
600
200
250
TABLE VII : Simulation data and temporal settings. Episode limits include demonstration-action warmup. Counts and horizons are in simulator control steps.
Behavior cloning (BC) is a practical offline imitation learning method, but it often fails when expert demonstrations are limited. Recent works have introduced a class of architectures named predictive inverse dynamics models (PIDMs) that combine a future-state predictor with an inverse dynamics model. While PIDMs often outperform BC, the reasons behind their benefits remain unclear. In this paper, we provide a theoretical explanation: PIDMs introduce a tradeoff. Conditioning the IDM on the predicted future state can significantly reduce variance, but the prediction itself introduces additional bias and variance. We establish conditions for PIDMs to achieve higher sample efficiency and lower prediction error than BC, with the gap widening when additional data sources are available. We validate the theoretical insights empirically in 2D navigation tasks, where BC requires up to five times (three times on average) more demonstrations than PIDM to reach comparable performance. Results are also illustrated in a complex 3D environment in a modern video game with high-dimensional visual inputs and stochastic transitions, where BC requires over 66% more samples than PIDM.
We introduce ABC, a fully open-source stack for manipulation with behavior cloning. At its core is ABC-130K: the largest open-source teleoperation dataset to date, featuring 3,500 hours of data spanning over 130K episodes across 195 diverse tasks. Furthermore, we open-source our accessible hardware setup, training infrastructure, and simulation pipeline. We also release 400 hours of sim-teleop data and provide a co-training recipe that produces correlated simulation and real-world evaluation, offering a reliable proxy for ablating model-design and training decisions before costly real-world evaluation. We explore various training recipes and compare common architectural choices for Diffusion Transformers (DiT) and Vision-Language-Action (VLA) models, grounding our findings in real-world evaluations. The resulting policies successfully execute dexterous tasks such as box folding and extracting credit cards from wallets. By providing a reproducible toolkit, we aim to place researchers on an equal footing, establishing the necessary foundation to learn the ABCs of Behavior Cloning together as a community.
Arthur Allshire, Himanshu Gaurav Singh, Ritvik Singh +15
Behavior-cloned policies often learn multiple behavior modes from demonstration datasets, including modes that are unsafe or otherwise undesired at deployment. For example, a policy trained on diverse handover demonstrations may learn to pass a knife blade-first. Standard remedies such as data curation and inference-time steering either require access to the original demonstrations for full retraining or add substantial inference-time overhead. To address this gap, we propose MoRE(Mode Redirection), which redirects policy rollouts toward desired behavior modes through a short "uncloning" step. Specifically, MoRE distills the redirection signal from a temporary mode classifier into the policy weights to steer behavior. A retain loss balances this edit by preserving desired-mode competence, allowing the standalone policy to suppress unwanted modes with zero inference-time overhead. Across eight simulated and real-world tasks, MoRE improves the average deployment success rate (SR) by 44 percentage points over the original mixed-mode policy. Among all compared adaptation and steering baselines, MoRE achieves the strongest SR and approaches the filtered-data retraining reference, while preserving task competence and inference speed. MoRE also generalizes across robot policy backbones, including Diffusion Policy and the Pi0.5 VLA, diverse task categories, and real-world deployments.
Hao Wang, Jiuzhou Lei, Dayou Li +5
Texas A&M University · University of Wisconsin · Northwestern University +1