Point It, Strike It: Direction-Conditioned Dynamic Manipulation of Deformable Linear Objects
Authors: Yi Yang, Xiang Fei, Lehong Wang, Zilin Dai, Ruogu Li, Jiting Cai, Liyao Chang, Xinyi Yang, +4 more
Organizations: The Robotics Institute, Carnegie Mellon University, 5000 Forbes Ave, Pittsburgh, PA 15213 USA. · School of Ocean and Civil Engineering, Shanghai Jiao Tong University, 800 Dongchuan Rd, Minhang District, Shanghai, China. · Zhiyuan College, Shanghai Jiao Tong University, 800 Dongchuan Rd, Minhang District, Shanghai, China. · John A. Paulson School of Engineering and Applied Sciences, Harvard University, 29 Oxford Street, Cambridge, MA 02138 USA.
Goal-conditioned dynamic manipulation of deformable linear objects has mainly specified goals as positions for a rope tip to reach. Many tasks, however, depend on how the tip arrives. We therefore study single-swing rope striking with goals that specify the tip's 3D position and arrival direction, across the workspace and on different ropes. This is challenging because rope dynamics are hard to model, no demonstrations exist, distinct swings reach the same goal with different reliability, and the sim-to-real gap extends beyond the rope. To address these challenges, we extend the state-of-the-art DLO simulator DeformX with GPU acceleration, a stable Cosserat rod solver, and a cross-flow aerodynamic model, yielding DeformX2.0, which is more than 20,000× faster. We then propose TRACE (Trace-rooted Adaptive Cross-Entropy), which generates striking data by warm-starting each new target from the stored swing whose tip path passes closest to it. Its cost penalizes rope bending and abrupt tip motion to favor repeatable swings. A conditional flow-matching policy trained on this data reaches 92.1% accuracy in simulation. Finally, we propose RECAP (Residual Calibration Policy), which fits the simulator's rope and rig parameters to a few calibration swings and adapts actions with a correction policy trained in simulation. On a real robot, across three ropes, RECAP raises success within 5cm from 72% to 87% for position goals, and within 10cm and 10° from 50% to 79% for goals that also specify the arrival direction.
Figures & tables
Fig. 2: System overview and network architectures. Left: TRACE generates goal–action pairs in GPU co-simulation to train the base policy. RECAP calibrates the real system once and corrects the nominal action for each goal before a single open-loop strike. (a) The base and correction policies use conditional flow matching (CFM), conditioned on g and (g,wˉ,Δ,η^) , respectively. (b) RECAP iteratively fits rope and rig parameters to measured calibration swings. These estimates condition the correction policy to adapt actions to the real system. Dashed lines indicate input.
Fig. 3: Experimental setup and simulation results. (a) Targets are selected within the workspace. (b) Arrival angle θ , measured after projecting the rope-tip velocity onto the tangent plane of the arm-centered sphere. (c) Radar-plot radius indicates data-generation coverage for each θ , and color indicates policy hit rate. (d,e) Data-generation coverage and policy accuracy, measured by solved generation targets and hit evaluation targets, respectively.
3D goal (position)
4D goal (position and direction)
Rope
Method
SR@5
SR@10
Distance
Best
Worst
SR@5
SR@10
Distance
Best
Worst
Angle
(%)
(%)
(cm)
(cm)
(cm)
(%)
(%)
(cm)
(cm)
(cm)
( ∘ )
Rope A
Base
68
91
4.8±4.9
1.3±1.1
15.4±5.9
23
41
13.4±12.0
1.4±0.5
38.7±25.4
9.8±17.3
Baseline
65
96
4.4±2.8
1.3±0.4
11.9±2.5
31
45
11.7±9.3
2.6±0.4
31.0±0.7
9.4±13.8
Ours
89
99
2.8±2.6
1.0±0.8
9.8±7.0
61
75
5.2±4.3
0.7±0.1
15.2±2.3
4.5±5.2
Rope B
Base
83
97
3.4±2.8
0.9±0.3
10.4±1.7
33
55
9.5±7.5
1.4±0.7
26.1±2.2
8.3±11.9
TABLE II: Hardware striking on three ropes, 25 targets × 3 swings per cell. Base: nominal simulator; Baseline: simulator identified by a learned network as in Wiggle&Go [ 6 ] , no correction; Ours: RECAP. SR@ d : success rate (%) within d of the target; 4D goals also require arriving within 10∘ of the commanded direction at ≥1 m/s along it. Distance and Angle (arrival-angle error): mean over all swings, with the std as a subscript. Best and Worst: the lowest and the highest per-target mean distance, with that target’s std. Mean is the mean over the three ropes. Best method per rope and goal type in bold.
Engine
Envs
Update rate (Hz)
Step time (ms)
Sim-to-wall time ratio
DeformX [ 2 ] (CPU, 1 thread)
1
12.9
77.5
0.22
Ours (RTX 4090)
1
136
7.4
2.3
Ours (RTX 4090)
1024
104
9.6
1780
Ours (RTX 4090)
2048
77.3
12.9
2640
Ours (RTX 4090)
8192
34.5
29.0
4710
TABLE I: Simulation throughput on the same 1 m, 20-segment rope, on one RTX 4090. Rates are per control step of the whole batch.
Fig. 4: Workspace maps on Rope A (real world) for 3D goals with Base, Baseline and Ours, and for 4D goals with Baseline and Ours: success rate (upper two rows) and mean tip–target distance (lower two rows) of each target over its three swings. A 3D goal succeeds within 5 cm; a 4D goal, which also requires arriving within 10∘ of the commanded direction at ≥1 m/s, is scored within 10 cm, as the column titles say. The distance scale saturates at 20 cm. Each quantity is shown in the front view (lateral x vs. height z from the wall-mounted base) and, directly below it on the same x axis, in the top view (lateral x vs. forward y ). Dots are the evaluated targets; the shaded field is a kernel-weighted interpolation between them, and the black star marks the wall-mounted robot base.
Method
Params
One-shot
Best-of-64
Miss
Angle
(M)
(%)
(%)
(cm)
( ∘ )
Ours
11.32
73.5
92.1
3.6
2.4
Naive FM
11.35
67.9
88.9
5.3
3.0
Regression
11.31
21.4
–
24.0
12.4
Nearest-neighbor
—
33.6
–
9.2
6.4
TABLE III: Learning-method ablation : We compare our model, naive flow matching, and our network with a regression head at matched capacity, alongside nearest-neighbor lookup in the training data. Miss and angle are means over the one-shot sample on n=5700 held-out goals. Regression falls below the lookup because it averages incompatible motion patterns into invalid swings.
Fig. 5: Learning multiple swings for one goal. We compare goal conditioning in every block (ours, left) with conditioning only at the input (naive FM, right). Stars mark two training swings wA,wB associated with goal g ; dots show generated samples. Shading shows a 2D slice of the conditional log-density logp(w∣g) , and arrows show the learned flow field at s=0.9 . (a,b) For swings from the same motion pattern, our model produces more concentrated samples. (c,d) For swings from different patterns, our model produces more clearly separated density peaks. Regression ( × ) predicts a single action that can average incompatible swings. (e) For the pair in (c,d), curves show the fraction of samples within a given distance of the nearer training swing, normalized by the pair’s separation in normalized action space.
Generator
Coverage (%) ↑
One-shot (%) ↑
Best-of-64 (%) ↑
Mean Lcurv↓
CEM w/ example
49.5
58.5
68.5
0.40
TRACE w/o example
96.8
75.5
83.9
0.51
TRACE
98.3
73.5
92.1
0.33
TABLE IV: Data-generation ablation evaluating workspace coverage, policy accuracy on 5700 random test goals, and trajectory quality measured by mean curvature loss. One-shot evaluates one policy sample; best-of-64 selects the candidate with the lowest TRACE loss. The example seeds higher-quality motion patterns, while TRACE improves coverage and accuracy of the dataset.