Manipulating a Rubik's Cube with a single dexterous hand is a challenging test of sustained, contact-rich control: the hand must execute successive layer turns while keeping the cube secure. Each turn requires some fingers to support the cube while others push a moving layer, release contact, and reset for the next move. To learn this coordination, we introduce FINGR (Future-supervised Interaction Network with Geometric Representations), a policy that combines finger-relative geometry with future interaction prediction. A shared point encoder expresses the cube relative to each fingertip and aggregates its points without depending on cubie indexing. Learned future tokens share the observation encoder and receive supervision for contact-force changes, layer-turn progress, and finger joint displacement at multiple time scales. The resulting representation conditions a flow policy that directly generates finger actions. On a real dexterous hand, our policy achieves 99.0% success over 300 turn attempts, compared with 79.7% for the base flow policy. Integrated with grasping and table-assisted regrasping, the policy solves all ten scrambled 2×2×2 cubes in a mean complete-system time of approximately 137 seconds. The project website is available at https://www.lyt0112.com/projects/FINGR
Figures & tables
Figure 1: Real cube solving through continuous layer turns. The four central frames show one 90∘ layer turn, from the starting grasp through finger reset. Repeating these learned motions, together with grasping and table-assisted regrasping, completes cube solving.
Figure 2: Policy architecture. Fingr uses a finger-relative cube position encoder to augment the five finger-state tokens. A transformer encoder processes these tokens together with tactile inputs, global cube geometry, and three future queries. At each future horizon, the future prediction head jointly predicts contact-force change, turn progress, and joint displacement during training. The full encoded memory, including the three future tokens, conditions the action head at deployment.
Figure 3: Layer-turn demonstrations. Each row shows six hand–cube configurations at approximately 18∘ increments across one 90∘ turn.
Method
U success ↑
L success ↑
Timeout ↓
Drop ↓
All success ↑
Time ↓
ACT
72.0 ± 8.7
83.3 ± 6.4
22.3 ± 6.5
0.0 ± 0.0
77.7 ± 6.5
6.39 ± 0.62
Diffusion Policy
78.0 ± 11.1
79.3 ± 13.0
20.3 ± 6.7
1.0 ± 0.0
78.7 ± 6.7
6.06 ± 0.14
Base Flow Policy
82.7 ± 13.0
76.7 ± 3.1
19.7 ± 7.5
0.7 ± 0.6
79.7 ± 8.0
5.67 ± 0.14
Fingr
99.3 ± 1.2
98.7 ± 1.2
1.0 ± 1.0
0.0 ± 0.0
99.0 ± 1.0
5.25 ± 0.20
Table 1: Real layer-turn comparison. Success, timeout, and drop entries are percentages. Time is the mean duration per attempt in seconds. Values are mean ± sample standard deviation over three rounds of 100 attempts. Bold marks column-wise best values, including ties.
Method
U success ↑
L success ↑
Timeout ↓
Drop ↓
All success ↑
Time ↓
Base Flow Policy
82.7 ± 13.0
76.7 ± 3.1
19.7 ± 7.5
0.7 ± 0.6
79.7 ± 8.0
5.67 ± 0.14
Local Geometry
92.7 ± 5.8
92.7 ± 1.2
7.0 ± 1.7
0.3 ± 0.6
92.7 ± 2.3
5.24 ± 0.04
Fingr
99.3 ± 1.2
98.7 ± 1.2
1.0 ± 1.0
0.0 ± 0.0
99.0 ± 1.0
5.25 ± 0.20
Table 2: Representation ablation. Local Geometry adds the finger-relative encoding to Base Flow. Fingr additionally includes future tokens with the joint contact, progress, and motion prediction objective.
Figure 4: Real layer turns. Six stages of each 90∘ turn show layer advancement and finger reset while the supporting fingers retain the cube.
Figure 5: Time spent in complete cube solving. Layer turns dominate execution; nine solves include a regrasp. Bars cover the full system time, detailed in Table 6 .
Scramble
# U
# L
U time ↓
L time ↓
Time/ 90∘↓
Regrasps
Plan ↓
Total ↓
1
8
10
6.30
3.98
5.01
1
1.56
140.87
2
9
8
9.92
5.34
7.76
1
1.97
193.59
3
7
5
5.69
6.26
5.93
1
1.55
119.48
4
10
9
5.28
4.26
4.79
1
1.83
141.62
5
8
9
3.35
4.07
3.73
0
0.50
84.08
6
8
8
4.50
4.21
4.36
1
3.95
134.14
Table 3: Complete physical cube solves. Counts and times cover planned 90∘ attempts. Plan includes motion planning and checks; Total spans initial grasp planning through final placement and robot return. Times are seconds.
Figure 6: From pickup to a solved cube. Left to right, top to bottom. Each pair of turn rows shows the first two 90∘ rotations after pickup or regrasp ( L then U ); later turns are omitted.
Appendix figures & tables4 assets
Supplementary material from the paper’s appendix.
Appendix
Method
Seed
Outcomes (%)
Mean time (s)
U success ↑
L success ↑
Timeout ↓
Drop ↓
All success ↑
U↓
L↓
All ↓
ACT
0
66
76
29
0
71
7.29
6.76
7.02
ACT
1
68
88
22
0
78
6.69
6.05
6.37
ACT
2
82
86
16
0
84
6.04
5.54
5.79
Diffusion Policy
0
80
92
13
1
86
6.40
5.42
5.91
Diffusion Policy
1
66
80
26
1
73
6.30
6.08
6.19
Appendix
Table 4: All layer-turn evaluation rounds. Outcome entries are percentages. U , L , and All times average every attempt of the corresponding direction or the full round, including failures, in seconds. Each row contains 50 U and 50 L attempts.
ID
Scramble
Planned solution sequences
1
R’ F U2 R2 U2 R’ F U’ F R2 U2
UL2U3L3 / LUL2ULU2L
2
R’ U2 F U F U’ F R2 F’ R U’
LU2LUL2 / U3LULULUL
3
U’ F’ R U2 R U’ R2 U’ F’ U2 R
LU3L / L2ULU3
4
U’ R’ U2 R’ F U’ F’ U2 R U’ R
LULU2L3UL / ULU2LULU2
5
U R’ F’ R2 U2 R2 U’ F U2 F’ R2
LULULULUL3U2LULU
6
U’ R’ U’ F’ U R’ U2 R U2 F U2
LU / U2LUL2U2LULUL2
Appendix
Table 5: Scramble and solution formulas. Solution formulas use grasp-relative positive U/L turns. The slash denotes a regrasp; an arrow denotes replanning under the same grasp. Small alignment corrections are timed separately.
ID
Pickup ↓
Layer turns ↓
Alignment ↓
Regrasp ↓
Place/return ↓
Observe ↓
Total ↓
1
8.72
90.20
1.80
24.95
11.40
3.80
140.87
2
8.85
132.00
0.00
39.05
10.61
3.08
193.59
3
9.02
71.10
0.70
24.56
10.30
3.79
119.48
4
8.06
91.10
0.00
25.85
12.96
3.66
141.62
5
8.31
63.40
0.00
0.00
10.10
2.26
84.08
6
9.24
69.70
0.00
40.43
11.38
3.40
134.14
Appendix
Table 6: Time breakdown for each complete solve. Seconds. Pickup includes initial planning, pregrasp, approach, closure, lift, and rotation into the manipulation pose. Regrasp includes intermediate placement and pickup in the new grasp. Observe includes the remaining state-observation and transition intervals.
Stage
Completed / attempted ↑
Success ↑
Duration ↓
Pregrasp and initial planning
10 / 10
100.0
3.33 ± 0.06
Approach
10 / 10
100.0
1.00 ± 0.00
Grasp closure
10 / 10
100.0
2.00 ± 0.05
Lift and manipulation-pose rotation
10 / 10
100.0
2.46 ± 0.41
Planned layer turn
181 / 185
97.8
4.59 ± 3.13
Alignment
4 / 4
100.0
2.00 ± 1.04
Appendix
Table 7: System-stage outcomes and durations. Success rates are percentages; time is mean ± sample standard deviation in seconds per operation. The first four stages form initial pickup.
Dexterous object rotation is a sequential contact problem: each support, release, and re-contact decision must both produce the desired object motion, and prepare the hand configuration for continued rotation. Existing reinforcement learning methods discover such movement patterns through trial and error on specific robotic hand embodiments, without explicitly accounting for how each contact transition affects the hand's ability to sustain object rotation in subsequent steps. We introduce DexMani, a framework that transfers human demonstrations as contact-conditioned manipulability evolution. This prior captures how successful human contact transitions reshape the object-rotation directions available to the hand. DexMani then learns this manipulability evolution and uses it to guide downstream reinforcement learning, enabling rotation skills to be acquired across robot embodiments with distinct kinematics and active-contact configurations. Across the Shadow Hand, Allegro Hand, and XHand, DexMani achieves the highest success rates in every evaluated setting for both seen and unseen objects. DexMani reaches an average success rate of 57.5% on LEAP Hand, outperforming other baselines and producing smoother rotatory motions. Project site: https://dexmani.github.io
Xiaoyang Chen, Shengcheng Luo, Haoran Guo +4
Shanghai Jiao Tong University · ShanghaiTech University · Beijing Institute for General Artificial Intelligence (BIGAI) +1
Recent work in humanoid whole-body control has found success with a simple recipe: retarget human motion to robot kinematic references, then train policies via reinforcement learning (RL) to track them. But how does this recipe transfer to dexterous manipulation? The answer is not obvious, as manipulation involves complex, contact-rich dynamics and requires delicate regulation of contact modes and forces. We present REGRIND, a minimalist retargeting-guided RL pipeline that learns dexterous manipulation policies from a single human demonstration. REGRIND retargets human hand-object motion to a robot reference that preserves hand-object spatial and contact relationships, trains a residual RL policy in simulation to track object-centric keypoints along that reference, and transfers the resulting policy zero-shot to hardware with careful system identification. The resulting policies produce fluid, human-like behavior on two different multi-fingered hands across contact-rich tool-use tasks, including operating a pair of scissors and turning a screwdriver. Through systematic hardware experiments, we identify and analyze the key factors that govern sim-to-real transfer in dexterous manipulation, offering practical guidance for retargeting-based learning in contact-rich settings. Videos and code are available at https://yunhaifeng.com/REGRIND.
Robotic foundation models (RFMs) are increasingly capable of general-purpose manipulation, yet reliable physical interaction remains challenging in contact-rich settings. We present FP2, a lightweight downstream interface that equips task-adapted RFMs with explicit force control while preserving their action-generation capability. FP2 adopts an action-regulation decomposition: the task-adapted RFM serves as a foundation policy responsible for task-level action generation, while a high-frequency force control policy focuses solely on interaction regulation. To condition force regulation on the ongoing manipulation, FP2 compresses foundation-policy contextual representations and combines them with wrench and proprioceptive histories to predict structured force-control parameters. We evaluate FP2 with four RFM backbones across four real-world contact-rich manipulation tasks. FP2 consistently improves task performance and force regulation quality over the corresponding foundation policies, while comparing favorably with representative force-aware and force-control baselines. Ablations further show that foundation-policy context and physical feedback are complementary for effective force regulation, while preserving foundation-policy action generation improves both efficiency and novel-object generalization. Project website: http://force-policy.github.io/fp2