Simulation-trained manipulation policies can exploit privileged state information to learn effective contact-rich behaviours, but deployment requires acting from partial observations such as noisy camera images. A common solution is teacher-student distillation, in which a visuomotor policy is trained to reproduce the actions of the privileged expert. This requires the student to jointly infer the task-relevant state and relearn the expert's action mapping that is already available. An alternative is to reuse the state-based expert and learn only a perceptual interface that reconstructs its missing state inputs. However, minimising the state estimate error alone does not necessarily minimise the downstream control error induced by these estimates. To bridge this gap, we train a visual state estimator using both direct state supervision and an action-consistency loss backpropagated through the frozen, differentiable expert. A scheduled objective first establishes a physically meaningful state estimate and progressively emphasises errors that affect the expert's actions. Across five goal-conditioned manipulation tasks, retaining the expert consistently outperforms direct pixel-to-action imitation from the same expert demonstration corpus. We further demonstrate sim-to-real transfer on a physical Panda robot, achieving 76% success without retraining the underlying expert.
Figures & tables
Fig. 2: The five environments: contact-rich manipulation tasks of varying difficulty, each with a single expert conditioned on a target object state.
Fig. 3: Closed-loop success rate per environment. Hatched: the frozen expert on the true state, repeated as a dashed reference. Orange: estimators with the retained expert (mixed-loss estimator with the linear ramp, Ls estimator, La -only estimator). Blue: action imitation (full-state flow policy and the best pixel BC baseline per task over flow, ACT, MIP and three vision trunks). Bars are means over three training seeds, error bars the sample s.d. The pixel-BC bars are flow on ResNet-18, seeded on all five tasks; every other pixel-BC cell remains single-seed. The wide PandaSphere La bar averages one converged and two collapsed seeds; protocol and values as in Tables II and III .
measured
estimated so
Task
dims
sq
object
goal-rel.
dima
PandaSphere
42
30
9
3
7
PandaCube
66
30
21
15
7
TraySpiral, TrayPlate
57
33
18
6
7
AllegroCube
96
60
21
15
16
TABLE I: Expert input per task. The estimator predicts only so , the object and goal-relative entries; sq is measured at deployment.
pixels → actions
state → actions
pixels → state → actions
Task
flow
ACT
MIP
full-state flow
Ls estimator
mixed-loss estimator
expert
PandaSphere
49.6 ± 1.3
43.5 R18
29.2 v3
85.0 ± 0.8
85.2 ± 0.3
94.8 ± 0.4
96.2
PandaCube
5.4 ± 1.1
6.1 R18
3.5 R18
45.7 ± 5.1
99.0 ± 0.4
99.2 ± 0.4
99.0
TraySpiral
36.4 ± 3.2
21.8 R18
32.0 v3
23.6 ± 1.6
90.0 ± 0.5
92.7 ± 0.8
96.8
TrayPlate
11.5 ± 1.5
10.7 R18
2.5 v3
4.4 ± 0.5
28.9 ± 1.7
64.1 ± 0.7
84.6
AllegroCube
0.7 ± 0.2
0.6 v3
0.5 v3
0.4 ± 0.1
79.2 ± 0.8
77.2 ± 0.9
82.6
TABLE II: Closed-loop success (%), protocol as in Sec. IV . ± is the sample s.d. over three training seeds; cells without it are single-seed and are the best of three trunks, the superscript naming the winner ( R18 ResNet-18, v2 DINOv2-LoRA, v3 DINOv3-LoRA); every trunk is listed in Table V . The flow column is ResNet-18 throughout, so it carries no superscript. Both estimators use DINOv3-LoRA, the mixed-loss estimator the linear ramp. Bold = best cell per row excluding the expert.
staged
ramped
flat
Task
expert
switch
linear
cosine
wa=1
wa=0.1
Ls only
La only
PandaSphere
96.2
94.6
94.8 ± 0.4
94.5 ± 0.3
93.6
92.9
85.2 ± 0.3
34.1 ± 52.4
PandaCube
99.0
99.5
99.2 ± 0.4
99.2 ± 0.3
98.4
99.4
99.0 ± 0.4
3.8 ± 2.0
TraySpiral
96.8
1.0
92.7 ± 0.8
94.3 ± 1.2
3.8
93.9
90.0 ± 0.5
0.4 ± 0.2
TrayPlate
84.6
0.8
64.1 ± 0.7
66.3 ± 1.8
0.8
62.8
28.9 ± 1.7
0.5 ± 0.2
AllegroCube
82.6
79.1
77.2 ± 0.9
78.1 ± 0.8
8.7
78.3
79.2 ± 0.8
0.0 ± 0.0
TABLE III: Mixed-loss schedules. Success rate (%), protocol as in Table II . expert : the frozen expert on the true state; Ls only : the state-only estimator, the baseline. Bold = best schedule per row, underline = below that baseline. The PandaSphere La -only entry averages one converged seed (94.6) with two collapsed ones (1.9, 5.9); see Sec. V-C .
Fig. 4: Domain-randomised training observations and real-world setup for PandaCube. Rows are the three cameras. Left: three simulated training episodes; middle: real-world observations; right: the physical setup.
Domain
Objective
Success (%)
Mean steps
Simulation
Mixed loss, linear ramp
99.2
26.1±12.5
Real world
Mixed loss, linear ramp
76.0
78.2±45.8
TABLE IV: Sim-to-real evaluation of the domain-randomised policy.
Task
Trunk
flow
ACT
MIP
PandaSphere
ResNet18
51.1
43.5
14.9
DINOv2-LoRA
20.7
26.0
10.8
DINOv3-LoRA
38.1
42.5
29.2
PandaCube
ResNet18
5.3
6.1
3.5
DINOv2-LoRA
2.4
1.6
0.7
DINOv3-LoRA
3.7
4.2
3.2
TABLE V: Pixel imitation across vision trunks and objectives. Closed-loop success (%), protocol as in Sec. IV , single training seed. Bold marks the best trunk within each task and objective; Table II quotes the bold ACT and MIP cells, and for flow the ResNet-18 trunk averaged over three training seeds.
MLP
full-state
Task
steps
0.18×
0.99×
0.2 – 0.3׆
4.1׆
flow
PandaSphere
40k
13.8
63.7 ± 1.9
23.7
63.4
85.0 ± 0.8
120k
56.6
82.3
57.6
66.2
93.3
PandaCube
40k
9.3
59.3
8.3
80.7
45.7 ± 5.1
120k
49.8
88.1
45.7
87.8
86.8
TraySpiral
40k
8.0
16.8
10.6
21.6
23.6 ± 1.6
TABLE VI: Closed-loop success (%) of a privileged-state MLP clone and the full-state flow policy at the shared 40k-step budget and at 120k steps; protocol as in Sec. IV . MLP widths relative to the frozen expert. ± = s.d. over three training seeds, other cells single-seed. † four-frame window and goal input, as for the imitation baselines; the other arms take one state frame. Bold = best per row.