Manipulating a Rubik's Cube with a single dexterous hand is a challenging test of sustained, contact-rich control: the hand must execute successive layer turns while keeping the cube secure. Each turn requires some fingers to support the cube while others push a moving layer, release contact, and reset for the next move. To learn this coordination, we introduce FINGR (Future-supervised Interaction Network with Geometric Representations), a policy that combines finger-relative geometry with future interaction prediction. A shared point encoder expresses the cube relative to each fingertip and aggregates its points without depending on cubie indexing. Learned future tokens share the observation encoder and receive supervision for contact-force changes, layer-turn progress, and finger joint displacement at multiple time scales. The resulting representation conditions a flow policy that directly generates finger actions. On a real dexterous hand, our policy achieves 99.0% success over 300 turn attempts, compared with 79.7% for the base flow policy. Integrated with grasping and table-assisted regrasping, the policy solves all ten scrambled 2×2×2 cubes in a mean complete-system time of approximately 137 seconds. The project website is available at https://www.lyt0112.com/projects/FINGR
Figures & tables
Figure 1: Real cube solving through continuous layer turns. The four central frames show one 90∘ layer turn, from the starting grasp through finger reset. Repeating these learned motions, together with grasping and table-assisted regrasping, completes cube solving.
Figure 2: Policy architecture. Fingr uses a finger-relative cube position encoder to augment the five finger-state tokens. A transformer encoder processes these tokens together with tactile inputs, global cube geometry, and three future queries. At each future horizon, the future prediction head jointly predicts contact-force change, turn progress, and joint displacement during training. The full encoded memory, including the three future tokens, conditions the action head at deployment.
Figure 3: Layer-turn demonstrations. Each row shows six hand–cube configurations at approximately 18∘ increments across one 90∘ turn.
Method
U success ↑
L success ↑
Timeout ↓
Drop ↓
All success ↑
Time ↓
ACT
72.0 ± 8.7
83.3 ± 6.4
22.3 ± 6.5
0.0 ± 0.0
77.7 ± 6.5
6.39 ± 0.62
Diffusion Policy
78.0 ± 11.1
79.3 ± 13.0
20.3 ± 6.7
1.0 ± 0.0
78.7 ± 6.7
6.06 ± 0.14
Base Flow Policy
82.7 ± 13.0
76.7 ± 3.1
19.7 ± 7.5
0.7 ± 0.6
79.7 ± 8.0
5.67 ± 0.14
Fingr
99.3 ± 1.2
98.7 ± 1.2
1.0 ± 1.0
0.0 ± 0.0
99.0 ± 1.0
5.25 ± 0.20
Table 1: Real layer-turn comparison. Success, timeout, and drop entries are percentages. Time is the mean duration per attempt in seconds. Values are mean ± sample standard deviation over three rounds of 100 attempts. Bold marks column-wise best values, including ties.
Method
U success ↑
L success ↑
Timeout ↓
Drop ↓
All success ↑
Time ↓
Base Flow Policy
82.7 ± 13.0
76.7 ± 3.1
19.7 ± 7.5
0.7 ± 0.6
79.7 ± 8.0
5.67 ± 0.14
Local Geometry
92.7 ± 5.8
92.7 ± 1.2
7.0 ± 1.7
0.3 ± 0.6
92.7 ± 2.3
5.24 ± 0.04
Fingr
99.3 ± 1.2
98.7 ± 1.2
1.0 ± 1.0
0.0 ± 0.0
99.0 ± 1.0
5.25 ± 0.20
Table 2: Representation ablation. Local Geometry adds the finger-relative encoding to Base Flow. Fingr additionally includes future tokens with the joint contact, progress, and motion prediction objective.
Figure 4: Real layer turns. Six stages of each 90∘ turn show layer advancement and finger reset while the supporting fingers retain the cube.
Figure 5: Time spent in complete cube solving. Layer turns dominate execution; nine solves include a regrasp. Bars cover the full system time, detailed in Table 6 .
Scramble
# U
# L
U time ↓
L time ↓
Time/ 90∘↓
Regrasps
Plan ↓
Total ↓
1
8
10
6.30
3.98
5.01
1
1.56
140.87
2
9
8
9.92
5.34
7.76
1
1.97
193.59
3
7
5
5.69
6.26
5.93
1
1.55
119.48
4
10
9
5.28
4.26
4.79
1
1.83
141.62
5
8
9
3.35
4.07
3.73
0
0.50
84.08
6
8
8
4.50
4.21
4.36
1
3.95
134.14
Table 3: Complete physical cube solves. Counts and times cover planned 90∘ attempts. Plan includes motion planning and checks; Total spans initial grasp planning through final placement and robot return. Times are seconds.
Figure 6: From pickup to a solved cube. Left to right, top to bottom. Each pair of turn rows shows the first two 90∘ rotations after pickup or regrasp ( L then U ); later turns are omitted.
Appendix figures & tables4 assets
Supplementary material from the paper’s appendix.
Appendix
Method
Seed
Outcomes (%)
Mean time (s)
U success ↑
L success ↑
Timeout ↓
Drop ↓
All success ↑
U↓
L↓
All ↓
ACT
0
66
76
29
0
71
7.29
6.76
7.02
ACT
1
68
88
22
0
78
6.69
6.05
6.37
ACT
2
82
86
16
0
84
6.04
5.54
5.79
Diffusion Policy
0
80
92
13
1
86
6.40
5.42
5.91
Diffusion Policy
1
66
80
26
1
73
6.30
6.08
6.19
Appendix
Table 4: All layer-turn evaluation rounds. Outcome entries are percentages. U , L , and All times average every attempt of the corresponding direction or the full round, including failures, in seconds. Each row contains 50 U and 50 L attempts.
ID
Scramble
Planned solution sequences
1
R’ F U2 R2 U2 R’ F U’ F R2 U2
UL2U3L3 / LUL2ULU2L
2
R’ U2 F U F U’ F R2 F’ R U’
LU2LUL2 / U3LULULUL
3
U’ F’ R U2 R U’ R2 U’ F’ U2 R
LU3L / L2ULU3
4
U’ R’ U2 R’ F U’ F’ U2 R U’ R
LULU2L3UL / ULU2LULU2
5
U R’ F’ R2 U2 R2 U’ F U2 F’ R2
LULULULUL3U2LULU
6
U’ R’ U’ F’ U R’ U2 R U2 F U2
LU / U2LUL2U2LULUL2
Appendix
Table 5: Scramble and solution formulas. Solution formulas use grasp-relative positive U/L turns. The slash denotes a regrasp; an arrow denotes replanning under the same grasp. Small alignment corrections are timed separately.
ID
Pickup ↓
Layer turns ↓
Alignment ↓
Regrasp ↓
Place/return ↓
Observe ↓
Total ↓
1
8.72
90.20
1.80
24.95
11.40
3.80
140.87
2
8.85
132.00
0.00
39.05
10.61
3.08
193.59
3
9.02
71.10
0.70
24.56
10.30
3.79
119.48
4
8.06
91.10
0.00
25.85
12.96
3.66
141.62
5
8.31
63.40
0.00
0.00
10.10
2.26
84.08
6
9.24
69.70
0.00
40.43
11.38
3.40
134.14
Appendix
Table 6: Time breakdown for each complete solve. Seconds. Pickup includes initial planning, pregrasp, approach, closure, lift, and rotation into the manipulation pose. Regrasp includes intermediate placement and pickup in the new grasp. Observe includes the remaining state-observation and transition intervals.
Stage
Completed / attempted ↑
Success ↑
Duration ↓
Pregrasp and initial planning
10 / 10
100.0
3.33 ± 0.06
Approach
10 / 10
100.0
1.00 ± 0.00
Grasp closure
10 / 10
100.0
2.00 ± 0.05
Lift and manipulation-pose rotation
10 / 10
100.0
2.46 ± 0.41
Planned layer turn
181 / 185
97.8
4.59 ± 3.13
Alignment
4 / 4
100.0
2.00 ± 1.04
Appendix
Table 7: System-stage outcomes and durations. Success rates are percentages; time is mean ± sample standard deviation in seconds per operation. The first four stages form initial pickup.