Retrieving objects from dense clutter requires rearrangement during which the manipulator can occlude objects while moving them. Repeated arm withdrawals to restore visibility interrupt execution. We introduce TRACE, a plan-conditioned imitation framework for retrieval under self-occlusion. A single unoccluded observation initializes a digital twin, where a privileged teacher generates a fixed nominal rollout. A recurrent student combines local rollout context, partial object observations, and proprioception to select actions that can correct deviations from the prediction. Behavior cloning initializes the student; DAgger refines it with teacher labels on student-visited states. The rollout remains fixed throughout execution, so the deployed student needs neither online teacher queries nor additional simulator rollouts during pushing. On 511 simulation test scenes, TRACE achieves 90.7% success versus 43.4% for nominal replay and 96.7% for the privileged closed-loop teacher. At a matched 26,373-label budget, student-state supervision achieves 87.8% versus 66.7% for expert-only cloning, demonstrating gains beyond additional labels. On a UR5e, TRACE achieves 90.0% success versus 95.0% for the closed-loop teacher, while reducing total execution time from 192.7 s to 67.3 s. It avoids the teacher's 16.8 sensing-related arm retractions per trial during pushing, retaining a final withdrawal for graspability evaluation. Code and data will be released at: https://trace-retrieval.github.io.
Figures & tables
Fig. 1: Hardware setup and plan-conditioned execution. (a) UR5e with fixed external RGB-D sensing. (b) TRACE execution during target occlusion; the nominal EEF path and active local plan window are shown for reference. (c) Final graspability evaluation. (d) Example pushing sequence through successful retrieval.
Fig. 2: End-effector motion primitives. The 16-action library contains four cardinal, four diagonal, and eight two-segment staircase motions, executed horizontal-first (H-XX) or vertical-first (V-XX).
Fig. 3: TRACE inference pipeline. (A) A clean initial scene estimate s^0 initializes a digital twin, where the frozen privileged teacher generates the fixed nominal rollout τˉ . (B) Current detections update object tracks; unavailable object geometry is zeroed while visibility, observation age, and target identity remain available. (C) The student combines the DeepSets scene representation, local plan context zt , end-effector features, and previous action. A GRU updates recurrent state ξt , and a categorical head predicts one of the 16 motion primitives. (D) The selected primitive is executed and the next partial observation closes the loop. After panel (A), online execution requires neither teacher queries nor additional simulator rollouts.
Fig. 4: Simulation setup and dataset. (a) Parallel Isaac Gym environments. (b)-(d) Representative training, validation, and test scenes (2,097/178/511); validation is used for model selection and the disjoint test set only for final evaluation.
Fig. 5: Qualitative real-robot comparison. Representative outcomes: Teacher Replay ends ungraspable, Spiral causes a workspace exit, PMBS reaches its 15-push limit, Online Teacher succeeds after repeated complete-scene reacquisition, and TRACE succeeds without online arm retraction. Timestamps are in seconds.
Method
Online state
Success
95% CI
OOW
Budget
(%)
(%)
(%)
Nominal-rollout methods
Teacher Replay
None
43.4
[39.1,48.0]
6.3
50.3
TRACE-BC
Partial
66.7
[63.7,69.8]
15.3
18.0
TRACE
Partial
90.7
[88.9,92.6]
7.2
2.1
Planning and heuristic baselines
TABLE I: Simulation comparison on the 511-scene test set. OOW denotes a workspace violation; Budget denotes the method-specific execution limit. Brackets are 95% stratified scene-bootstrap CIs from 2,000 resamples.
Training data
Labels
Success
Δ
(%)
(pp, 95% CI)
DAgger aggregation
Expert BC
26,373
66.7
–
DAgger R1
57,125
87.1
+20.4[17.5,23.2]
DAgger R2
84,429
89.0
+2.0[−0.1,4.0]
TRACE (DAgger R3)
111,555
90.7
+1.7[−0.2,3.5]
TABLE II: Effect of student-state supervision on the 511-scene test set. Results average three seeds; all fits use 10,000 updates. Δ is relative to the preceding round in the aggregation block and Expert BC in the matched-budget blocks; brackets are paired 95% CIs.
Variant
Success
Δ vs. TRACE
OOW
Budget
Steps
(%)
(pp, 95% CI)
(%)
(%)
TRACE ( K=4 )
90.7
—
7.2
2.1
14.8
w/o GRU
85.8
−4.9[−6.5,−2.4]
8.0
6.2
15.0
K=1
88.5
−2.2[−3.7,+0.2]
8.4
3.1
14.6
K=2
89.9
−0.8[−2.1,+1.4]
7.8
2.3
14.7
K=8
89.3
−1.4[−2.7,+0.8]
7.7
3.0
14.7
TABLE III: Plan-context horizon and recurrent-memory ablations on the 511-scene test set (three seeds). Δ and 95% CIs are paired against TRACE ( K=4 ); Budget is the teacher-relative step/travel limit. Steps is the mean over successful episodes.
Method
Execution
Online information
Success (%)
OOW (%)
Init./plan (s)
Online exec. (s)
Grasp+lift (s)
Total (s)
Arm retractions
Teacher Replay
Open loop
Initial plan only
47.5
0.0
27.5
11.6
10.4
49.5
0.0
Spiral
Closed loop
Target pose
77.5
17.5
6.3
29.2
8.9
44.4
2.8
PMBS
Closed loop
Complete scene
85.0
10.0
124.9
7.4
8.8
141.1
3.9
Online Teacher
Closed loop
Complete scene
95.0
0.0
10.4
173.3
9.0
192.7
16.8
TRACE (Ours)
Closed loop
Partial scene + τˉ
90.0
0.0
30.4
24.8
12.1
67.3
0.0
TABLE IV: Real-robot evaluation on 20 scenes with two trials each. OOW denotes a workspace violation. Arm retractions count visual-state reacquisition; TRACE ’s final graspability-check retraction is excluded. Other unsuccessful trials reached the method-specific execution limit.
Retrieving objects buried beneath clutter is both challenging and time-consuming, as complex support relationships make manipulation particularly difficult. Existing methods either focus on support relations and rely on sequential grasping to remove occluding objects, or perform preparatory actions such as pushing to facilitate subsequent grasps. However, these approaches are often inefficient and treat physical interactions as isolated auxiliary steps. In this paper, we propose RetrDex, an efficient framework for dexterous arm-hand systems to learn object retrieval in cluttered scenes. Our approach leverages large-scale parallel reinforcement learning (RL) in diverse cluttered scenes and incorporates a spatially aware representation that encodes occlusion patterns and spatial relationships among the target, the dexterous hand, and surrounding clutter. This representation enables the policy to develop diverse manipulation skills (e.g., pushing, stirring, and poking) that actively clear occluders. We evaluate RetrDex on 16 household objects across varied clutter configurations, and obtain strong retrieval performance and efficiency on both seen and unseen targets. Furthermore, we demonstrate successful zero-shot transfer to a real-world dexterous multi-fingered robot system, validating the practical applicability of our method. Videos can be found on our project website: https://RetrDex.github.io.
Fengshuo Bai, Yu Li, Jie Chu +5
Shanghai Jiao Tong University · PKU-PsiBot Joint Lab · Peking University +2
Parametric imitation learning via behavior cloning can suffer from poor generalization to out-of-distribution states due to compounding errors during deployment. We show that reusing the training data during inference via a semi-parametric retrieval-based imitation learning approach can alleviate this challenge. We present Difference-Aware Retrieval Policies for Imitation Learning (DARP), a semi-parametric retrieval-based imitation learning approach that addresses this limitation by reparameterizing the imitation learning problem in terms of local neighborhood structure rather than direct state-to-action mappings. Instead of learning a global policy, DARP trains a model to predict actions based on k-nearest neighbors from expert demonstrations, their corresponding actions, and the relative distance vectors between neighbor states and query states. DARP requires no additional assumptions beyond those made for standard behavior cloning -- it does not require additional data collection, online expert feedback, or task-specific knowledge. We demonstrate consistent performance improvements of 15-46% over standard behavior cloning across diverse domains, including continuous control and robotic manipulation, and across different representations, including high-dimensional visual features. Code and demos are available at https://weirdlabuw.github.io/darp-site/.
Quinn Pfeifer, Ethan Pronovost, Paarth Shah +3
1Paul G. Allen School of Computer Science & Engineering, University of Washington · 2Toyota Research Institute · 3Google DeepMind +1
Robots under autonomous operation may require decisions based on evidence that is no longer visible. We study delayed-evidence tasks, where an early cue disappears before a later decision point, so visually similar observations can require different actions. In these settings, the current observation is not a sufficient state for control. We introduce TRAjectory-routed Causal Evidence (TRACE), a memory framework for visuomotor imitation policies. TRACE stores task-relevant visual and robot-state evidence, such as object identity, target choice, or route-dependent state, in a fixed-size latent memory that remains bounded over long episodes. Instead of indexing memory by raw time or manually provided task labels, TRACE uses path signatures: compact, order-sensitive features of the executed robot-state trajectory. These signatures do not store the visual cue itself; rather, they provide trajectory-conditioned keys for writing and retrieving the evidence stored when the cue was visible. When the robot later reaches an ambiguous observation, the policy conditions on TRACE memory to recover the missing context and choose the correct branch. TRACE attaches through lightweight adapters to policies, without changing the policy backbone, action head, or imitation objective. Across real-world long-horizon manipulation tasks with visually ambiguous branch points, TRACE improves branch selection and task success over alternative baselines, including short-history and recurrent memory. Project page: https://jeong-zju.github.io/trace
Zihao Li, Ranpeng Qiu, Yincong Chen +2
School of Computer Science and the Australian Centre for Robotics at the University of Sydney.