Pose error in precision assembly can corrupt not only what a robot observes but also the coordinate frame in which it acts. On the FORGE benchmark, the official state-based policy succeeds in 97% to 99% of episodes with the true pose but only 32% to 60% at the benchmark's σ=5 mm pose-noise setting. The same estimated pose enters the observation and anchors the action frame, making the offset unidentifiable from proprioceptive state alone before contact. We supply this missing information during training in two ways. A privileged teacher observes the offset in simulation, while clean demonstrations can instead be relabelled into the displaced frame in closed form. The deployed student is trained with behaviour cloning followed by one DAgger round and receives only noisy state, a raw wrench window, and two RGB cameras at test time. On the unmodified FORGE tasks, the teacher-route student maintains 92% to 99% success across σ=0 to 5 mm, while six non-privileged baselines fall to 2% to 80%. A matched behaviour-cloning experiment isolates the source of this robustness. With the same student architecture, data budget, and training procedure, demonstrations generated without offset access yield only 31.5% success at σ=5 mm, whereas privileged and relabelled demonstrations reach 88.7% and 95.8%. Deployed zero-shot on a Franka, the student reaches 83.3% pooled success at σ=5 mm against 34.4% for the strongest state-based policy. Deployable sensing alone is insufficient. Robustness requires supervision that encodes compensation for the latent frame offset.Project page: https://drychang.github.io/displaced-frame/
Figures & tables
Fig. 1: The displaced action frame. Top: ϵ contaminates both the observation and action frame, making it unidentifiable from proprioceptive state before contact. Bottom: (a) initial and goal states of the three tasks. (b) Pose error to scale. At 5 mm, the displaced frame lies outside the peg hole.
Fig. 2: (a) Two ways to encode offset compensation in demonstrations: a privileged teacher or closed-form relabelling, followed by BC and one DAgger round. (b) The 26.5 M-parameter student uses only deployable state, wrench, and RGB sensing from a wrist and a third-person (TP) camera, fused into a single-shot 7 -dim action at 15 Hz. (c) The same policy and sensing are deployed unchanged on the Franka.
task
model
0 mm
1 mm
2.5 mm
5 mm
peg
T-A official baseline
98.7 / 98.2
97.3 / 94.7
73.7 / 74.5
31.8 / 33.3
T-A+ noise-augmented
84.1 / 77.3
78.0 / 75.1
58.7 / 56.2
29.9 / 25.8
T-B privileged (oracle)
99.3 / 96.6
99.3 / 96.7
99.7 / 97.5
99.3 / 96.6
student (ours)
98.3 / 96.2
98.6 / 96.2
98.3 / 95.4
96.6 / 91.8
gear
T-A official baseline
98.6 / 100.0
97.5 / 99.2
86.9 / 90.6
49.9 / 53.8
T-A+ noise-augmented
94.9 / 99.7
94.7 / 99.7
92.1 / 99.3
77.9 / 90.0
TABLE I: Success rate (%) across the noise spectrum, averaged over three seeds ( n=256 /cell). Cells report dynamics randomisation on/off. T-A+ retrains T-A at σ=5 mm, and T-B is the privileged oracle.
demonstrations
0 mm
2.5 mm
5 mm
10 mm
T-A under noise (3 seeds)
98.6 ± 0.8
70.6 ± 1.9
31.5 ± 2.2
8.7 ± 0.5
priv. T-B (3 seeds)
95.6 ± 3.3
93.9 ± 6.3
88.7 ± 7.1
70.7 ± 9.7
clean, relabelled (3 seeds)
98.0 ± 0.8
96.5 ± 1.8
95.8 ± 1.0
84.2 ± 2.6
clean, no shift (s0)
98.0
63.7
26.2
7.8
clean, state shift only (s0)
98.4
98.8
94.1
66.8
clean, relabelled + DAgger (s0)
98.8
99.6
100.0
94.5
TABLE II: Matched peg ablation ( n=256 /cell). The first three rows differ only in demonstration construction and report three-seed means ± sd. Controls and + DAgger use seed 0.
Fig. 3: Where the weak arm presses (peg, 128 paired episodes per arm and noise level, identical offsets). Left: at σ=5 mm the non-privileged arm first touches a median 6.2 mm off the axis and its failures stay on the diagonal (median 7.7 mm at depth), whereas the relabelled arm touches at 4.3 mm and is on-axis at depth ( 1.4 mm). Right: outcomes at σ=2.5 and 5 mm. All but one of the 122 non-inserted episodes never entered the hole.
Fig. 4: Success across the noise spectrum ( n=256 /cell, three-seed means within the training range). Beyond the dashed 5 mm training boundary, T-A, T-A+, and T-B are single-seed, while the student is single-seed on peg beyond 10 mm. Right: the teacher swap of Table II , with all arms BC-only except the + DAgger curve (three-seed means). The non-privileged arm tracks its teacher’s degradation, whereas relabelled demonstrations match the privileged-teacher result. The state shift explains most of the recovery, no shift reproduces the collapse, and DAgger brings the relabelled peg arm to the oracle level.
method
family
peg
gear
nut
T-B
privileged oracle (upper bound)
99.7/99.3
98.2/98.2
99.9/99.2
SRSA [ 7 ]
skill retrieval
0 3.1/ 0 2.3
13.7/13.3
22.3/17.2
C3
RMA-style adaptation [ 6 ]
59.4/30.1
69.5/46.5
76.2/54.3
ARCH-INSERT [ 8 ]
hierarchical RL
64.5/29.7
89.5/53.9
88.7/60.5
T-A ( ϵ critic)
asymmetric critic [ 18 ]
73.1/35.6
96.5/79.7
91.0/71.9
T-A ( U law)
noise-aug., U[0,5] law
74.6/31.6
91.0/74.2
91.8/67.6
TABLE III: Success rate (%) at 2.5/5 mm with dynamics randomisation on under one frozen protocol. Our method, T-A, and T-A+ use three seeds. External baselines are single-seed re-implementations: SRSA is outside its released skill library, and C3 is an RMA-style adapter rather than CoRMA itself. The matched- U[0,5] and ϵ -critic T-A variants are also single-seed controls.
task
full
no force
no vision
state only
peg
98.3/98.3/96.6 †
86.3/82.2/78.1 †
47.5 † /50.0/30.7 †
35.5/30.5/19.1
gear
98.8/99.3/98.2 †
98.8/98.0/98.2 †
91.8/82.0/63.7
68.4/49.2/34.8
nut
99.1/99.0/96.5 †
92.2/91.2/91.7 †
76.6/68.4/52.0
49.2/50.0/32.0
TABLE IV: Modality ablation, retrained per condition (%, 0/2.5/5 mm, dynamics randomisation on). Cells marked † report three seeds (no-force SD: peg 10 to 15 , gear ≤2.6 , nut 7 to 9 ). All others are single-seed and noted in Sec. VIII .
Fig. 5: The three tasks at their initial and goal states, in simulation (top, fixed-camera view) and on the Franka (bottom). Part colours and lighting differ between simulation and the printed setup, with photometric augmentation reducing the appearance gap. Pose offsets are injected into the taught fixture pose, while the physical parts remain fixed across trials.
task
policy
0 mm
1 mm
2.5 mm
5 mm
peg
T-A+
86.7
83.3
63.3
20.0
student
90.0
93.3
86.7
80.0
gear
T-A+
90.0
90.0
73.3
36.7
student
93.3
90.0
83.3
86.7
nut
T-A+
90.0
83.3
70.0
46.7
student
93.3
90.0
86.7
83.3
TABLE V: Real-robot success rate (%, 30 trials per cell) on a Franka Panda with printed parts and a dark work surface. Offsets of the stated σ enter both the observation and action frame, matching simulation. T-A+ is the seed-0 noise-augmented state baseline without cameras. Pooled rows combine all three tasks ( n=90 ).
As end-to-end robotic policies are progressively deployed in the real world to solve real tasks, they face a gap between the training and inference conditions. Scaling the amount and diversity of the training data has shown some success in improving zero-shot generalization, yet robots still fail when faced with new, unseen test conditions. For instance, while robots with fixed frames of reference are common, those with moving frames pose a greater challenge for deployment. To address this specific instance of the issue, we present a study of strategies for encoding the robot's proprioceptive state to improve both in- and out-of-distribution performance at test time. Through a systematic study of joint representations, we find that a simple episode-wise relative frame provides the best trade-off between task performance and robustness, outperforming the baselines in extensive real-robot experiments conducted in a realistic test environment. The results suggest a practical path to leveraging data collected by robots with varying frames of reference and deployment to unseen test configurations.
Maxime Alvarez, Ryo Watanabe, Paul Crook +4
TELEXISTENCE Inc, Foundation Model Division, Japan. · The University of Tokyo, Japan.
When using reinforcement learning (RL) for contact-rich robotic manipulation, vision can provide task-relevant information that complements robot proprioception. However, vision-enabled policies tend to overfit to the visual conditions seen during training, limiting their robustness and transferability. We present a human-in-the-loop RL framework that employs teacher-student distillation to achieve robust performance across multiple task variants, trained entirely in the real world without requiring domain randomization or data augmentation. A vision-enabled teacher distills its knowledge into a vision-free student that relies solely on pose, twist, and wrench sensing, combining fast training with strong task generalization. On the real-world NIST assembly benchmark board, our approach achieves 95% overall success after approximately 50 minutes of training on 3 representative tasks, including robust generalization to 8 unseen task variants. Fine-tuning with distillation achieves full success on the most challenging task. We demonstrate that the resulting policies outperform baselines in both robustness and adaptability. Page: https://tuwien-asl.github.io/VE2VF/.
Victor Kowalski, Chengxi Li, Dongheui Lee
Autonomous Systems, Technische Universitaet Wien (TU Wien), Vienna, Austria · Institute of Robotics and Mechatronics (DLR), German Aerospace Center, Wessling, Germany
Robots inspecting an assembly must determine which parts are present and whether they are correctly installed. During egocentric assembly assistance, head motion and workpiece handling reveal evidence for these checks, while spoken state confirmations link observations to procedural outcomes. We introduce INSPECT, which learns robot view preferences from records of a smart-glasses assistant that answers part queries and provides next-step guidance. Presence-Invariant TwinSwap (PI-TwinSwap) calibrates object evidence through paired identity interventions. Claim-indexed supervision separates evidence requirements from camera-reproducible observation changes. Object-centered calibration adapts relative view preferences to robot poses, while clause-level screening checks predicted evidence. The robot selects views using only its current observation and known poses, without candidate images. Evaluation uses annotated assistant-video replay to simulate state feedback, without target-domain view labels for policy training. On images of physical gearbox assemblies, INSPECT achieves the highest view utility among the compared non-oracle policies and raises human-rated full verifiability from 34.8% to 41.7% compared with keeping the current view. On commercial angle-grinder recordings in IMPACT, the transferred relative-view selector increases the correct decision rate from 50.6% to 54.3% with a frozen perception head. The source code is available at https://github.com/Kratos-Wen/INSPECT.
Di Wen, Kailun Yang, Wenhao Guo +7
Karlsruhe Institute of Technology, Karlsruhe 76131, Germany · Hunan University, Changsha 410012, China · Independent Researcher