We present FoldBack, a self-correcting masked generative policy for long-horizon garment folding. Existing long-trajectory policies may continue after a missed or slipped grasp even when the garment has not reached the intended configuration. We structure FoldBack's recovery mechanisms around three inference-time decisions: when to refine and verify, how to roll back, and where and how to retry. FoldBack aligns refinement and grasp verification with pick-and-place events, returns the robot to a retryable pre-grasp configuration while preserving successful grasps, and selectively regenerates the failed segment and selected future actions while avoiding previous failed grasp locations. To our knowledge, FoldBack is the first editable full-trajectory policy to unify these decisions, enabling failed interactions to be detected, undone, and repaired before execution continues, without recovery demonstrations or base-policy retraining. Across 33 real garments from six categories, FoldBack achieves 75.2% final folding success and 0.837 final-mask IoU, versus 45.7% and 0.689 for the strongest prior baseline.
Figures & tables
Fig. 1: Failure handling in garment folding. Unlike baseline policies that may continue or stall after a failed grasp, FoldBack verifies the grasp outcome, rolls back, repairs the trajectory, and retries before advancing.
Fig. 2: Overview of FoldBack. Key-Action-Aligned Inference aligns trajectory refinement with pick-and-place events; grasp monitoring and verified rollback determine whether a grasp-critical window is added to the verified trajectory history or kept editable for recovery; structured plan repair then remasks and resamples the editable trajectory after failure.
Fig. 3: Repair under an asymmetric bimanual grasp failure.
Fig. 4: Real-world FoldBack recovery. (a) Short-sleeved T-shirt: recovery from repeated grasp failures during multi-stage folding. (b) Long-sleeved T-shirt: recovery from an asymmetric bimanual failure while preserving the successful grasp.
Method
Napkin
Towel
Pants
SC
RS
FS
IoU
SC
RS
FS
IoU
SC
RS
FS
IoU
DP3
71.7
–
56.7
0.717
60.5
–
53.3
0.703
66.7
–
60.0
0.706
SKIL
75.0
–
63.3
0.755
65.8
–
60.0
0.743
60.0
–
53.3
0.731
MGP-Long
78.3
–
70.0
0.786
73.7
–
66.7
0.791
73.3
–
66.7
0.764
FoldBack w/o R&R
83.3
–
73.3
0.828
81.6
–
73.3
0.815
80.0
–
66.7
0.786
DP3 + VR
85.0
53.8
76.7
0.823
78.9
60.0
73.3
0.814
76.7
55.6
66.7
0.788
TABLE I: Real-world long-horizon garment-folding performance. ‘VR’, ‘R&R’, and ‘SR’ denote Verified Rollback, Rollback-and-Repair, and Structured Repair, respectively.
One-Corner Slip
Two-Corner Slip
Method
Det. ↑
Pres. ↑
RS ↑
FS ↑
Det. ↑
RS ↑
FS ↑
DP3
N/A
N/A
N/A
6.7
N/A
N/A
0.0
DP3 + VR
86.7
38.5
53.8
33.3
86.7
46.2
20.0
MGP-Long
N/A
N/A
N/A
13.3
N/A
N/A
6.7
FB w/o Rollback
93.3
92.9
57.1
40.0
93.3
42.9
13.3
FB w/o SR + FF
93.3
64.3
50.0
26.7
100.0
40.0
20.0
TABLE II: Recovery from controlled grasp-slip perturbations on towel folding. ‘FB’, ‘VR’, ‘SR’, and ‘FF’ denote FoldBack, verified rollback, structured repair, and fixed failed window.
Initial Orientation (Pants)
Unseen Instance
Non-Canonical
Method
Seen
Mod.
Large
Nap.
Tow.
Pants
S-T
L-T
S-S
Avg.
Nap.
Tow.
DP3
60.0/0.706
46.7/0.661
33.3/0.603
60.0/0.704
40.0/0.642
50.0/0.671
10.0/0.486
0.0/0.421
10.0/0.458
28.3/0.564
40.0/0.648
33.3/0.612
MGP-Long
66.7/0.764
60.0/0.724
40.0/0.642
70.0/0.774
50.0/0.710
60.0/0.731
20.0/0.596
10.0/0.521
10.0/0.530
36.7/0.644
53.3/0.704
46.7/0.674
FoldBack (ours)
80.0/0.851
73.3/0.817
60.0/0.752
90.0/0.863
80.0/0.833
70.0/0.812
60.0/0.780
50.0/0.756
50.0/0.742
66.7/0.798
80.0/0.826
73.3/0.798
TABLE III: Generalization performance, reported as final folding success rate (%) / final-mask IoU. S-T, L-T, and S-S denote short-sleeve T-shirt, long-sleeve T-shirt, and short-sleeve shirt, respectively.
Due to the highly deformable nature of garments, training a generalizable policy for robotic T-shirt folding and unfolding remains a significant challenge. In this work, we present a large-scale synthetic dataset for robotic T-shirt folding and unfolding, covering 6 robotic embodiments, 1K T-shirts, 1K environmental assets, and 120K episodes with rich annotations, which can be used to train a wide range of manipulation policies. We first follow the FoldNet pipeline to generate a large-scale dataset of physically simulatable T-shirts with diverse appearances and annotated semantic keypoints. Based on these semantic keypoints, we then generate manipulation demonstrations for different robotic embodiments through a unified rule-based framework. We use these demonstrations to train visuomotor policies, and experimental results demonstrate that models trained solely on our synthetic data can achieve over 90% end-to-end task success rates when directly deployed to unseen real-world environments and previously unseen T-shirts from arbitrary initial configurations. Project URL: https://pku-epic.github.io/FoldNetXX/.
Robotic garment unfolding is essential for downstream tasks, yet quasi-static methods require repeated actions, while existing dynamic approaches predominantly rely on bimanual flinging. We present RotateIt!, a single-arm framework that uses adaptive axial rotation for dynamic garment unfolding. To the best of our knowledge, it is the first unfolding framework to employ dynamic axial rotation as its primary manipulation primitive. From a randomly initialized tabletop configuration, the robot selects a rotation-effective grasp and rotates the lifted garment about an approximately fixed anchor, generating inertial tension that separates overlapping layers within a compact workspace. A grasp ranker selects the anchor, while an online residual policy adapts the rotation extent and speed, thereby determining the release timing. Across seen and unseen simulated garments and eight unseen real garments, RotateIt! improves success within three attempts by 44.0-61.0 percentage points over quasi-static pick-and-place. The simulation-trained policies transfer zero-shot to the real world, achieving 75.6% success, 41% higher first-attempt coverage, and 26% higher final coverage. The resulting states further enable autonomous robotic folding without manual rearrangement.
Zeqing Zhang, Zuokun Xie, Ao Fang +4
Nanyang Technological University · The University of Hong Kong · Chinese Academy of Sciences
I describe my solution to the LeHome Challenge 2026, an ICRA 2026 competition on bimanual garment folding. The system placed 1st of 62 teams in the online (simulation) round and 2nd in the real-world final. It improves a vision-language-action (VLA) policy with a reinforcement-learning loop. The policy is its own value function: the same network that predicts actions also predicts success, progress, and a few task-relevant future quantities, and those predictions drive advantage estimation, live failure detection, and candidate selection. The work mostly recombines existing RL ideas with engineering and optimization contributions that can be used together as one recipe or individually: AWR + RECAP combined for flow-matching VLA; an asynchronous distributed training / rollout pipeline through HuggingFace Hub; inference-time hyperparameters optimization via Thompson sampling; a sim-to-real recipe with camera-alignment tooling, heavy augmentation and DAgger-like HIL data collection.