Pose error in precision assembly can corrupt not only what a robot observes but also the coordinate frame in which it acts. On the FORGE benchmark, the official state-based policy succeeds in 97% to 99% of episodes with the true pose but only 32% to 60% at the benchmark's σ=5 mm pose-noise setting. The same estimated pose enters the observation and anchors the action frame, making the offset unidentifiable from proprioceptive state alone before contact. We supply this missing information during training in two ways. A privileged teacher observes the offset in simulation, while clean demonstrations can instead be relabelled into the displaced frame in closed form. The deployed student is trained with behaviour cloning followed by one DAgger round and receives only noisy state, a raw wrench window, and two RGB cameras at test time. On the unmodified FORGE tasks, the teacher-route student maintains 92% to 99% success across σ=0 to 5 mm, while six non-privileged baselines fall to 2% to 80%. A matched behaviour-cloning experiment isolates the source of this robustness. With the same student architecture, data budget, and training procedure, demonstrations generated without offset access yield only 31.5% success at σ=5 mm, whereas privileged and relabelled demonstrations reach 88.7% and 95.8%. Deployed zero-shot on a Franka, the student reaches 83.3% pooled success at σ=5 mm against 34.4% for the strongest state-based policy. Deployable sensing alone is insufficient. Robustness requires supervision that encodes compensation for the latent frame offset.Project page: https://drychang.github.io/displaced-frame/
Figures & tables
Fig. 1: The displaced action frame. Top: ϵ contaminates both the observation and action frame, making it unidentifiable from proprioceptive state before contact. Bottom: (a) initial and goal states of the three tasks. (b) Pose error to scale. At 5 mm, the displaced frame lies outside the peg hole.
Fig. 2: (a) Two ways to encode offset compensation in demonstrations: a privileged teacher or closed-form relabelling, followed by BC and one DAgger round. (b) The 26.5 M-parameter student uses only deployable state, wrench, and RGB sensing from a wrist and a third-person (TP) camera, fused into a single-shot 7 -dim action at 15 Hz. (c) The same policy and sensing are deployed unchanged on the Franka.
task
model
0 mm
1 mm
2.5 mm
5 mm
peg
T-A official baseline
98.7 / 98.2
97.3 / 94.7
73.7 / 74.5
31.8 / 33.3
T-A+ noise-augmented
84.1 / 77.3
78.0 / 75.1
58.7 / 56.2
29.9 / 25.8
T-B privileged (oracle)
99.3 / 96.6
99.3 / 96.7
99.7 / 97.5
99.3 / 96.6
student (ours)
98.3 / 96.2
98.6 / 96.2
98.3 / 95.4
96.6 / 91.8
gear
T-A official baseline
98.6 / 100.0
97.5 / 99.2
86.9 / 90.6
49.9 / 53.8
T-A+ noise-augmented
94.9 / 99.7
94.7 / 99.7
92.1 / 99.3
77.9 / 90.0
TABLE I: Success rate (%) across the noise spectrum, averaged over three seeds ( n=256 /cell). Cells report dynamics randomisation on/off. T-A+ retrains T-A at σ=5 mm, and T-B is the privileged oracle.
demonstrations
0 mm
2.5 mm
5 mm
10 mm
T-A under noise (3 seeds)
98.6 ± 0.8
70.6 ± 1.9
31.5 ± 2.2
8.7 ± 0.5
priv. T-B (3 seeds)
95.6 ± 3.3
93.9 ± 6.3
88.7 ± 7.1
70.7 ± 9.7
clean, relabelled (3 seeds)
98.0 ± 0.8
96.5 ± 1.8
95.8 ± 1.0
84.2 ± 2.6
clean, no shift (s0)
98.0
63.7
26.2
7.8
clean, state shift only (s0)
98.4
98.8
94.1
66.8
clean, relabelled + DAgger (s0)
98.8
99.6
100.0
94.5
TABLE II: Matched peg ablation ( n=256 /cell). The first three rows differ only in demonstration construction and report three-seed means ± sd. Controls and + DAgger use seed 0.
Fig. 3: Where the weak arm presses (peg, 128 paired episodes per arm and noise level, identical offsets). Left: at σ=5 mm the non-privileged arm first touches a median 6.2 mm off the axis and its failures stay on the diagonal (median 7.7 mm at depth), whereas the relabelled arm touches at 4.3 mm and is on-axis at depth ( 1.4 mm). Right: outcomes at σ=2.5 and 5 mm. All but one of the 122 non-inserted episodes never entered the hole.
Fig. 4: Success across the noise spectrum ( n=256 /cell, three-seed means within the training range). Beyond the dashed 5 mm training boundary, T-A, T-A+, and T-B are single-seed, while the student is single-seed on peg beyond 10 mm. Right: the teacher swap of Table II , with all arms BC-only except the + DAgger curve (three-seed means). The non-privileged arm tracks its teacher’s degradation, whereas relabelled demonstrations match the privileged-teacher result. The state shift explains most of the recovery, no shift reproduces the collapse, and DAgger brings the relabelled peg arm to the oracle level.
method
family
peg
gear
nut
T-B
privileged oracle (upper bound)
99.7/99.3
98.2/98.2
99.9/99.2
SRSA [ 7 ]
skill retrieval
0 3.1/ 0 2.3
13.7/13.3
22.3/17.2
C3
RMA-style adaptation [ 6 ]
59.4/30.1
69.5/46.5
76.2/54.3
ARCH-INSERT [ 8 ]
hierarchical RL
64.5/29.7
89.5/53.9
88.7/60.5
T-A ( ϵ critic)
asymmetric critic [ 18 ]
73.1/35.6
96.5/79.7
91.0/71.9
T-A ( U law)
noise-aug., U[0,5] law
74.6/31.6
91.0/74.2
91.8/67.6
TABLE III: Success rate (%) at 2.5/5 mm with dynamics randomisation on under one frozen protocol. Our method, T-A, and T-A+ use three seeds. External baselines are single-seed re-implementations: SRSA is outside its released skill library, and C3 is an RMA-style adapter rather than CoRMA itself. The matched- U[0,5] and ϵ -critic T-A variants are also single-seed controls.
task
full
no force
no vision
state only
peg
98.3/98.3/96.6 †
86.3/82.2/78.1 †
47.5 † /50.0/30.7 †
35.5/30.5/19.1
gear
98.8/99.3/98.2 †
98.8/98.0/98.2 †
91.8/82.0/63.7
68.4/49.2/34.8
nut
99.1/99.0/96.5 †
92.2/91.2/91.7 †
76.6/68.4/52.0
49.2/50.0/32.0
TABLE IV: Modality ablation, retrained per condition (%, 0/2.5/5 mm, dynamics randomisation on). Cells marked † report three seeds (no-force SD: peg 10 to 15 , gear ≤2.6 , nut 7 to 9 ). All others are single-seed and noted in Sec. VIII .
Fig. 5: The three tasks at their initial and goal states, in simulation (top, fixed-camera view) and on the Franka (bottom). Part colours and lighting differ between simulation and the printed setup, with photometric augmentation reducing the appearance gap. Pose offsets are injected into the taught fixture pose, while the physical parts remain fixed across trials.
task
policy
0 mm
1 mm
2.5 mm
5 mm
peg
T-A+
86.7
83.3
63.3
20.0
student
90.0
93.3
86.7
80.0
gear
T-A+
90.0
90.0
73.3
36.7
student
93.3
90.0
83.3
86.7
nut
T-A+
90.0
83.3
70.0
46.7
student
93.3
90.0
86.7
83.3
TABLE V: Real-robot success rate (%, 30 trials per cell) on a Franka Panda with printed parts and a dark work surface. Offsets of the stated σ enter both the observation and action frame, matching simulation. T-A+ is the seed-0 noise-augmented state baseline without cameras. Pooled rows combine all three tasks ( n=90 ).
Autonomous Systems, Technische Universitaet Wien (TU Wien), Vienna, Austria · Institute of Robotics and Mechatronics (DLR), German Aerospace Center, Wessling, Germany