Simulation-trained manipulation policies can exploit privileged state information to learn effective contact-rich behaviours, but deployment requires acting from partial observations such as noisy camera images. A common solution is teacher-student distillation, in which a visuomotor policy is trained to reproduce the actions of the privileged expert. This requires the student to jointly infer the task-relevant state and relearn the expert's action mapping that is already available. An alternative is to reuse the state-based expert and learn only a perceptual interface that reconstructs its missing state inputs. However, minimising the state estimate error alone does not necessarily minimise the downstream control error induced by these estimates. To bridge this gap, we train a visual state estimator using both direct state supervision and an action-consistency loss backpropagated through the frozen, differentiable expert. A scheduled objective first establishes a physically meaningful state estimate and progressively emphasises errors that affect the expert's actions. Across five goal-conditioned manipulation tasks, retaining the expert consistently outperforms direct pixel-to-action imitation from the same expert demonstration corpus. We further demonstrate sim-to-real transfer on a physical Panda robot, achieving 76% success without retraining the underlying expert.
Figures & tables
Fig. 2: The five environments: contact-rich manipulation tasks of varying difficulty, each with a single expert conditioned on a target object state.
Fig. 3: Closed-loop success rate per environment. Hatched: the frozen expert on the true state, repeated as a dashed reference. Orange: estimators with the retained expert (mixed-loss estimator with the linear ramp, Ls estimator, La -only estimator). Blue: action imitation (full-state flow policy and the best pixel BC baseline per task over flow, ACT, MIP and three vision trunks). Bars are means over three training seeds, error bars the sample s.d. The pixel-BC bars are flow on ResNet-18, seeded on all five tasks; every other pixel-BC cell remains single-seed. The wide PandaSphere La bar averages one converged and two collapsed seeds; protocol and values as in Tables II and III .
measured
estimated so
Task
dims
sq
object
goal-rel.
dima
PandaSphere
42
30
9
3
7
PandaCube
66
30
21
15
7
TraySpiral, TrayPlate
57
33
18
6
7
AllegroCube
96
60
21
15
16
TABLE I: Expert input per task. The estimator predicts only so , the object and goal-relative entries; sq is measured at deployment.
pixels → actions
state → actions
pixels → state → actions
Task
flow
ACT
MIP
full-state flow
Ls estimator
mixed-loss estimator
expert
PandaSphere
49.6 ± 1.3
43.5 R18
29.2 v3
85.0 ± 0.8
85.2 ± 0.3
94.8 ± 0.4
96.2
PandaCube
5.4 ± 1.1
6.1 R18
3.5 R18
45.7 ± 5.1
99.0 ± 0.4
99.2 ± 0.4
99.0
TraySpiral
36.4 ± 3.2
21.8 R18
32.0 v3
23.6 ± 1.6
90.0 ± 0.5
92.7 ± 0.8
96.8
TrayPlate
11.5 ± 1.5
10.7 R18
2.5 v3
4.4 ± 0.5
28.9 ± 1.7
64.1 ± 0.7
84.6
AllegroCube
0.7 ± 0.2
0.6 v3
0.5 v3
0.4 ± 0.1
79.2 ± 0.8
77.2 ± 0.9
82.6
TABLE II: Closed-loop success (%), protocol as in Sec. IV . ± is the sample s.d. over three training seeds; cells without it are single-seed and are the best of three trunks, the superscript naming the winner ( R18 ResNet-18, v2 DINOv2-LoRA, v3 DINOv3-LoRA); every trunk is listed in Table V . The flow column is ResNet-18 throughout, so it carries no superscript. Both estimators use DINOv3-LoRA, the mixed-loss estimator the linear ramp. Bold = best cell per row excluding the expert.
staged
ramped
flat
Task
expert
switch
linear
cosine
wa=1
wa=0.1
Ls only
La only
PandaSphere
96.2
94.6
94.8 ± 0.4
94.5 ± 0.3
93.6
92.9
85.2 ± 0.3
34.1 ± 52.4
PandaCube
99.0
99.5
99.2 ± 0.4
99.2 ± 0.3
98.4
99.4
99.0 ± 0.4
3.8 ± 2.0
TraySpiral
96.8
1.0
92.7 ± 0.8
94.3 ± 1.2
3.8
93.9
90.0 ± 0.5
0.4 ± 0.2
TrayPlate
84.6
0.8
64.1 ± 0.7
66.3 ± 1.8
0.8
62.8
28.9 ± 1.7
0.5 ± 0.2
AllegroCube
82.6
79.1
77.2 ± 0.9
78.1 ± 0.8
8.7
78.3
79.2 ± 0.8
0.0 ± 0.0
TABLE III: Mixed-loss schedules. Success rate (%), protocol as in Table II . expert : the frozen expert on the true state; Ls only : the state-only estimator, the baseline. Bold = best schedule per row, underline = below that baseline. The PandaSphere La -only entry averages one converged seed (94.6) with two collapsed ones (1.9, 5.9); see Sec. V-C .
Fig. 4: Domain-randomised training observations and real-world setup for PandaCube. Rows are the three cameras. Left: three simulated training episodes; middle: real-world observations; right: the physical setup.
Domain
Objective
Success (%)
Mean steps
Simulation
Mixed loss, linear ramp
99.2
26.1±12.5
Real world
Mixed loss, linear ramp
76.0
78.2±45.8
TABLE IV: Sim-to-real evaluation of the domain-randomised policy.
Task
Trunk
flow
ACT
MIP
PandaSphere
ResNet18
51.1
43.5
14.9
DINOv2-LoRA
20.7
26.0
10.8
DINOv3-LoRA
38.1
42.5
29.2
PandaCube
ResNet18
5.3
6.1
3.5
DINOv2-LoRA
2.4
1.6
0.7
DINOv3-LoRA
3.7
4.2
3.2
TABLE V: Pixel imitation across vision trunks and objectives. Closed-loop success (%), protocol as in Sec. IV , single training seed. Bold marks the best trunk within each task and objective; Table II quotes the bold ACT and MIP cells, and for flow the ResNet-18 trunk averaged over three training seeds.
MLP
full-state
Task
steps
0.18×
0.99×
0.2 – 0.3׆
4.1׆
flow
PandaSphere
40k
13.8
63.7 ± 1.9
23.7
63.4
85.0 ± 0.8
120k
56.6
82.3
57.6
66.2
93.3
PandaCube
40k
9.3
59.3
8.3
80.7
45.7 ± 5.1
120k
49.8
88.1
45.7
87.8
86.8
TraySpiral
40k
8.0
16.8
10.6
21.6
23.6 ± 1.6
TABLE VI: Closed-loop success (%) of a privileged-state MLP clone and the full-state flow policy at the shared 40k-step budget and at 120k steps; protocol as in Sec. IV . MLP widths relative to the frozen expert. ± = s.d. over three training seeds, other cells single-seed. † four-frame window and goal input, as for the imitation baselines; the other arms take one state frame. Bold = best per row.
Vision-based imitation learning has enabled impressive robotic manipulation skills, but action imitation alone provides limited supervision of the geometric consequences of robot behavior. To address this limitation, we introduce Implicit Scene Supervision (ISS) Policy, a 3D visuomotor diffusion policy with a DiT backbone that predicts continuous action sequences from point-cloud observations. ISS augments action diffusion with a supervised robot motion predictor that maps generated actions and robot-state context to end-effector motion, and then uses the predicted motion together with gripper intent to forecast future point-cloud representations. By explicitly modeling the intermediate transition from action to robot motion, ISS encourages the policy to capture how its actions affect the surrounding 3D scene. We further introduce asymmetric gradient routing to separate direct motion regression from scene-level policy supervision, together with a change-balanced objective that accounts for variations in scene-change magnitude. These auxiliary objectives provide dynamics-aware geometric supervision using only expert demonstrations, without requiring additional annotations or auxiliary modules at inference time. ISS Policy achieves state-of-the-art performance on single-arm manipulation tasks in MetaWorld and dexterous manipulation tasks in Adroit, while real-world dual-arm experiments further demonstrate its effectiveness on physical robotic manipulation. The resulting framework preserves the scalable DiT backbone and standard diffusion-policy control interface. Code and videos will be released.
Recent visual imitation learning systems have widely adopted multi-camera setups with wrist-mounted cameras as the de facto standard. However, manipulation from a single global view remains challenging, as the policy should capture fine-grained interaction details and identify task-relevant regions without local wrist views. To address this challenge, we present Spatially Conditioned Diffusion Policy (SCDP), a diffusion-based visuomotor policy that achieves precise and robust manipulation in a single-camera setting. Our key idea is that end-effector trajectories can serve as visual attention anchors that reflect task-relevant regions. Building on this idea, SCDP consists of two key components: (i) a visual encoder that produces multi-scale feature maps to capture both broader context and fine-grained visual features, and (ii) a spatial conditioning module that samples point-wise features along intermediate end-effector trajectories in the diffusion loop. Extensive simulation experiments show that SCDP consistently outperforms strong single-view baselines and achieves performance comparable to multi-camera baselines. Real-world experiments further demonstrate precise manipulation and robustness to visual distractors, highlighting the potential of single-camera imitation learning.
Seoyoon Kim, Kanghyun Kim, Dongwoo Ko +2
Korea Advanced Institute of Science and Technology (KAIST) · Neuromeka
We present GHOST, a framework for learning visuomotor manipulation policies that generalize beyond the training distribution. GHOST factorizes control into (i) a high-level policy that predicts the next sub-goal as a distribution over 3D end-effector poses from multi-view RGB-D observations, and (ii) a low-level goal-conditioned controller that executes embodiment-specific actions. To condition image-based policies on 3D goals, we introduce a simple spatial interface that projects predicted goals into the image plane and represents them as end-effector heatmaps. Across a suite of manipulation tasks, this hierarchical factorization consistently improves performance and robustness compared to a flat Diffusion Policy. Further, we show that this hierarchical interface also makes it easy to incorporate human demonstrations without relying on (noisy) action retargeting. As sub-goals are largely embodiment-agnostic, we train the high-level policy on human video to specify how learned skills should be applied and composed, while keeping the low-level policy trained purely on robot data. This hierarchy enables adaptation to novel objects and task variations using a small number of human demonstrations.
Sriram Krishna, Ben Eisner, Haotian Zhan +5
Robotics Institute, Carnegie Mellon University · UMass Amherst