Retrieving objects from dense clutter requires rearrangement during which the manipulator can occlude objects while moving them. Repeated arm withdrawals to restore visibility interrupt execution. We introduce TRACE, a plan-conditioned imitation framework for retrieval under self-occlusion. A single unoccluded observation initializes a digital twin, where a privileged teacher generates a fixed nominal rollout. A recurrent student combines local rollout context, partial object observations, and proprioception to select actions that can correct deviations from the prediction. Behavior cloning initializes the student; DAgger refines it with teacher labels on student-visited states. The rollout remains fixed throughout execution, so the deployed student needs neither online teacher queries nor additional simulator rollouts during pushing. On 511 simulation test scenes, TRACE achieves 90.7% success versus 43.4% for nominal replay and 96.7% for the privileged closed-loop teacher. At a matched 26,373-label budget, student-state supervision achieves 87.8% versus 66.7% for expert-only cloning, demonstrating gains beyond additional labels. On a UR5e, TRACE achieves 90.0% success versus 95.0% for the closed-loop teacher, while reducing total execution time from 192.7 s to 67.3 s. It avoids the teacher's 16.8 sensing-related arm retractions per trial during pushing, retaining a final withdrawal for graspability evaluation. Code and data will be released at: https://trace-retrieval.github.io.
Figures & tables
Fig. 1: Hardware setup and plan-conditioned execution. (a) UR5e with fixed external RGB-D sensing. (b) TRACE execution during target occlusion; the nominal EEF path and active local plan window are shown for reference. (c) Final graspability evaluation. (d) Example pushing sequence through successful retrieval.
Fig. 2: End-effector motion primitives. The 16-action library contains four cardinal, four diagonal, and eight two-segment staircase motions, executed horizontal-first (H-XX) or vertical-first (V-XX).
Fig. 3: TRACE inference pipeline. (A) A clean initial scene estimate s^0 initializes a digital twin, where the frozen privileged teacher generates the fixed nominal rollout τˉ . (B) Current detections update object tracks; unavailable object geometry is zeroed while visibility, observation age, and target identity remain available. (C) The student combines the DeepSets scene representation, local plan context zt , end-effector features, and previous action. A GRU updates recurrent state ξt , and a categorical head predicts one of the 16 motion primitives. (D) The selected primitive is executed and the next partial observation closes the loop. After panel (A), online execution requires neither teacher queries nor additional simulator rollouts.
Fig. 4: Simulation setup and dataset. (a) Parallel Isaac Gym environments. (b)-(d) Representative training, validation, and test scenes (2,097/178/511); validation is used for model selection and the disjoint test set only for final evaluation.
Fig. 5: Qualitative real-robot comparison. Representative outcomes: Teacher Replay ends ungraspable, Spiral causes a workspace exit, PMBS reaches its 15-push limit, Online Teacher succeeds after repeated complete-scene reacquisition, and TRACE succeeds without online arm retraction. Timestamps are in seconds.
Method
Online state
Success
95% CI
OOW
Budget
(%)
(%)
(%)
Nominal-rollout methods
Teacher Replay
None
43.4
[39.1,48.0]
6.3
50.3
TRACE-BC
Partial
66.7
[63.7,69.8]
15.3
18.0
TRACE
Partial
90.7
[88.9,92.6]
7.2
2.1
Planning and heuristic baselines
TABLE I: Simulation comparison on the 511-scene test set. OOW denotes a workspace violation; Budget denotes the method-specific execution limit. Brackets are 95% stratified scene-bootstrap CIs from 2,000 resamples.
Training data
Labels
Success
Δ
(%)
(pp, 95% CI)
DAgger aggregation
Expert BC
26,373
66.7
–
DAgger R1
57,125
87.1
+20.4[17.5,23.2]
DAgger R2
84,429
89.0
+2.0[−0.1,4.0]
TRACE (DAgger R3)
111,555
90.7
+1.7[−0.2,3.5]
TABLE II: Effect of student-state supervision on the 511-scene test set. Results average three seeds; all fits use 10,000 updates. Δ is relative to the preceding round in the aggregation block and Expert BC in the matched-budget blocks; brackets are paired 95% CIs.
Variant
Success
Δ vs. TRACE
OOW
Budget
Steps
(%)
(pp, 95% CI)
(%)
(%)
TRACE ( K=4 )
90.7
—
7.2
2.1
14.8
w/o GRU
85.8
−4.9[−6.5,−2.4]
8.0
6.2
15.0
K=1
88.5
−2.2[−3.7,+0.2]
8.4
3.1
14.6
K=2
89.9
−0.8[−2.1,+1.4]
7.8
2.3
14.7
K=8
89.3
−1.4[−2.7,+0.8]
7.7
3.0
14.7
TABLE III: Plan-context horizon and recurrent-memory ablations on the 511-scene test set (three seeds). Δ and 95% CIs are paired against TRACE ( K=4 ); Budget is the teacher-relative step/travel limit. Steps is the mean over successful episodes.
Method
Execution
Online information
Success (%)
OOW (%)
Init./plan (s)
Online exec. (s)
Grasp+lift (s)
Total (s)
Arm retractions
Teacher Replay
Open loop
Initial plan only
47.5
0.0
27.5
11.6
10.4
49.5
0.0
Spiral
Closed loop
Target pose
77.5
17.5
6.3
29.2
8.9
44.4
2.8
PMBS
Closed loop
Complete scene
85.0
10.0
124.9
7.4
8.8
141.1
3.9
Online Teacher
Closed loop
Complete scene
95.0
0.0
10.4
173.3
9.0
192.7
16.8
TRACE (Ours)
Closed loop
Partial scene + τˉ
90.0
0.0
30.4
24.8
12.1
67.3
0.0
TABLE IV: Real-robot evaluation on 20 scenes with two trials each. OOW denotes a workspace violation. Arm retractions count visual-state reacquisition; TRACE ’s final graspability-check retraction is excluded. Other unsuccessful trials reached the method-specific execution limit.