Wrench-ACT: Enhancing Robot Policies for Contact Rich Behavior Using Direct Wrench Control
Authors: Johannes Hechtl, Yannik Blei, Simon Ball, Reihaneh Mirjalili, Michael Krawez, Seongjin Bien, Philipp Schmitt, Wolfram Burgard
Organizations: Siemens Research and Predevelopment · Department of Computer Science and Artificial Intelligence, University of Technology Nuremberg, Germany
While contact-rich manipulation requires deliberate regulation of interaction forces, recent approaches to robot manipulation learning predominantly represent actions as target positions or poses. Even methods that incorporate force sensing either use it solely as an observation or, when predicting forces as part of the output, rely on a hybrid force controller. In this paper, we propose an imitation learning policy that predicts wrenches as its sole action output for direct use by a pure force controller. Our studies suggest that force-domain imitation learning depends critically on data collection, with force-feedback teleoperation improving policy performance by capturing the operator's deliberate force regulation. Using Action Chunking with Transformers (ACT) as the base architecture, we train single-task models on bilateral wrench demonstrations and evaluate them on five contact-rich manipulation tasks. The wrench policy matches or outperforms position-based baselines across all tasks, with gains varying according to the degree of deliberate force regulation each task requires. Cross-condition ablations show that the bilateral data collection interface and the wrench action space each contribute independently to performance. To support further research, we will release over 1000 wrench-action demonstrations spanning these tasks on a companion website upon publication.
Figures & tables
Fig. 1 : Method overview of the proposed pure wrench approach. (A) Data collection via bilateral teleoperation. Operator wrench commands are executed by the follower arm and recorded, while the follower position is mirrored back to the leader. (B) Inference. Multimodal observations (RGB streams and robot state) are processed by the policy to directly predict a 6D target wrench, entirely bypassing intermediate positional control for contact-rich manipulation.
Fig. 2 : Overview of the five contact-rich manipulation tasks evaluated in this work (a - e) and the ablation evaluating robustness to box-height variation (f)
Fig. 3 : Bilateral Teleoperation Setup. Operator wrench commands, measured by the leader (left), are executed by the follower arm (right). The follower position is mirrored back to the leader.
Fig. 4 : Results of the contact force analysis. Bilateral teleoperation with a wrench action space ( πbw ) achieves high success rates while maintaining relatively low median 99th-percentile force magnitude and a narrow distribution of applied contact forces.
Action Space
Wrench
Position
Dataset
Bilateral
πbw
πbp
VR data
πvrw
πvrp
TABLE I : The four policies resulting from the combination of data collection interface and action representation.
Task
πbw
πbp
πvrw
πvrp
Peg Insertion
74%
2%
2%
6%
Fuse Clipping
88%
72%
52%
96%
Fan Insertion
58%
28%
0%
8%
Industrial Connector
80%
26%
4%
12%
Pen Writing
84%
46%
86%
78%
Average
76.8 %
34.8%
28.8%
40%
TABLE II : Success rates (%) across tasks and policies. πbw : bilateral data with wrench actions; πbp : bilateral data with pose actions; πvrw : VR data with wrench actions; πvrp : VR data with pose actions. All policies are evaluated on 50 rollouts for all tasks.
Fig. 5 : Ablation on the inference frequency on the industrial connector task. We show the dependence of success rate on inference speed. Contrary to a widespread assumption, an inference rate of 50Hz is sufficient for a success rate of 80% in our setting.
Task
πbw
πbp
πvrw
πvrp
Pen Writing, Nominal Height
84 %
46%
86%
78%
Pen Writing, Box Raised
76%
8%
56%
24%
TABLE III : Success rates (%) for Pen Writing at the nominal box height and with the box raised by 2.5 cm.