Visuomotor policies for humanoid loco-manipulation must generalize across object poses and lighting from only a handful of demonstrations. 3D Diffusion Policy (DP3) conditions a diffusion-based action generator on point-cloud features, but its PointNet-style encoder has no built-in equivariance to the rotations, translations, and scalings (SIM(3)) that manipulation tasks respect. EquiBot closed this gap for wheeled manipulators with a SIM(3)-equivariant Vector Neuron Network (VNN) encoder. We extend this to a substantially more complex embodiment, the 43-joint Unitree G1 humanoid, and propose EquivDP3: a two-stage policy where a high-level diffusion planner with a SIM(3)-equivariant VNN encoder emits 6 Hz whole-body command chunks, executed at 50 Hz by a frozen, pre-trained RL locomotion policy and a differential inverse-kinematics module for the arms, trained end-to-end by behavior cloning. Across two simulated IsaacLab benchmarks and four non-equivariant baselines (5-100 demonstrations, in- and out-of-distribution), EquivDP3's advantage concentrates in the low-data regime: at 5-10 demonstrations it reaches 67.1% success versus 38.2-52.4% for the baselines, while by 50-100 all encoders converge (74.3-85.2%) and the ordering is no longer meaningful. A proprioception-only control confirms this gap is genuinely perceptual: with the point cloud removed, success drops to 31% vs. 60% (EquivDP3) at 5 demonstrations and 78% vs. 99% at 10, but vanishes by 50-100, showing the high-data plateau reflects a benchmark ceiling, not five encoders learning the same invariance. The encoder costs only 0.8 ms of extra latency per action chunk over the PointNet encoder it replaces. Baking geometric symmetry into a hierarchical diffusion policy's perception backbone is a practical, nearly free way to improve data efficiency for humanoid loco-manipulation when demonstrations are scarce.
Figures & tables
Fig. 1 : EquivDP3 two-stage architecture. A diffusion planner with a SIM(3)-invariant point-cloud encoder runs at ∼ 6 Hz and emits whole-body command chunks. A low-level controller at 50 Hz splits each command between an RL-trained locomotion policy (velocity/angular/height/torso → leg joints) and an inverse-kinematics module (wrist pose → arm joints); both feed a joint-space PD controller.
Encoder
Rotation
Translation
Scale
EquivDP3 (ours)
3.9×10−7
2.7×10−7
3.8×10−7
PointNet
2.3×10−1
5.3×10−1
1.5×10−1
Ours, no LayerNorm
4.1×10−7
6.0×10−7
5.2
TABLE I : Empirical invariance check. Relative error ∥f(g⋅P)−f(P)∥/∥f(P)∥ over random transforms; 10−7 is float32 round-off. The final row shows that disabling the encoder’s LayerNorms removes scale invariance only.
Fig. 2 : Overview of the two evaluation tasks. Task 1 (Galileo Loco-Manipulation, left) requires full whole-body coordination across four subgoals: approach/walk to the shelf, bimanual pick, rotate & carry, and place into the sorting bin. Task 2 (G1 Pick-and-Place, right) fixes the base and requires only two subgoals: bimanual pick and place into the target bin.
Task
Subgoals
# subgoals
Galileo Loco-Manip.
approach/walk to shelf → bimanual pick → rotate & carry → place into bin
4
G1 Pick-and-Place
pick → place (fixed base)
2
TABLE II : Evaluation tasks.
Fig. 3 : Overall success rate vs. number of demonstrations, for Task 1 (Galileo Loco-Manipulation, left column) and Task 2 (G1 Pick-and-Place, right column), under in-distribution (top row) and out-of-distribution (bottom row) conditions. EquivDP3 (teal) leads in aggregate at 5–10 demonstrations; by 50–100 all encoders converge, and it is not the top performer in every individual setting.
EquivDP3
PointNet
PointNet+Aug
DGCNN
PCT
Task
ndemo
ID
OOD
ID
OOD
ID
OOD
ID
OOD
ID
OOD
Task 1
5
79.00 ± 7.91
47.00 ± 9.60
43.00 ± 9.53
49.00 ± 9.61
73.00 ± 8.58
40.00 ± 9.43
17.00 ± 7.33
20.00 ± 7.77
4.00 ± 4.14
2.00 ± 3.23
10
89.00 ± 6.19
60.00 ± 9.43
53.00 ± 9.60
37.50 ± 6.65
81.00 ± 7.63
59.00 ± 9.47
4.00 ± 4.14
12.00 ± 6.41
89.00 ± 6.19
66.00 ± 9.13
50
97.00 ± 3.71
65.00 ± 9.19
94.00 ± 4.85
61.50 ± 6.68
98.00 ± 3.23
78.00 ± 8.03
94.00 ± 4.85
64.00 ± 9.25
75.00 ± 5.96
46.00 ± 9.59
100
98.00 ± 3.23
86.00 ± 6.81
97.00 ± 3.71
70.00 ± 6.30
99.00 ± 2.64
84.00 ± 7.16
100.00 ± 1.85
71.00 ± 8.76
100.00 ± 1.85
78.00 ± 8.03
Task 2
5
60.00 ± 9.43
27.00 ± 8.58
10.00 ± 4.19
17.00 ± 7.33
34.00 ± 6.51
21.00 ± 7.91
85.00 ± 6.99
66.00 ± 9.13
42.33 ± 5.56
16.67 ± 5.95
TABLE III : Success rate (%) by encoder, task, and demonstration count, with in-distribution (ID) and out-of-distribution (OOD) as paired sub-columns under each encoder. Each cell reports the point estimate and the half-width of its 95% Wilson score interval [ 33 ] ( p±h ), computed directly from the logged per-episode outcomes. Episodes per cell: 63 cells use 100, 6 use 150, 9 use 200, 2 use 300. Cells evaluated more than once on the identical checkpoint are pooled, with n the pooled count; which cells were repeated reflects development history, not selection on outcome.
ndemo=5
ndemo=10
Encoder
SG1
SG2
SG3
SG4
SG1
SG2
SG3
SG4
EquivDP3
100
90
87
79
100
98
97
89
PointNet
100
32
27
43
100
79
53
53
PointNet+Aug
100
77
76
73
100
99
98
81
DGCNN
100
42
29
17
100
100
100
4
PCT
100
18
17
4
100
99
96
89
TABLE IV : Task 1 subgoal event rates (%), ID, at the two low-data budgets. SG1 approach, SG2 lift, SG3 carry to bin (gated on SG2), SG4 task success (not gated on SG2).
Fig. 4 : ID → OOD success-rate drop (percentage points) for both tasks, at (a) ndemo=100 and (b) ndemo=10 . EquivDP3 shows the smallest drop among all encoders at n=100 on Task 1, and DGCNN nearly collapses on Task 1 at n=10 (both ID and OOD performance near floor).
n=5
n=10
n=50
n=100
In-distribution
Blind (no point cloud) a
31
78
92
94
EquivDP3 (ours)
60
99
90
96
Gap (EquivDP3 − Blind)
+29
+21
−2
+2
Out-of-distribution
Blind (no point cloud) a
21
62
65
64
TABLE V : Proprioception-only control on Task 2 across the full demonstration-budget grid of Table III , under both distribution conditions. Bold gaps are significant by Fisher’s exact test ( p<0.05 ).
Fig. 5 : Left: median per-chunk inference latency over all single-environment evaluation runs, well under the 160 ms real-time budget (dashed line) implied by an 8-step chunk at 50 Hz; EquivDP3 costs 0.8 ms more than PointNet, and PCT is the slowest. Right: mean episode length ( ± std) in simulated steps at ndemo=100 , ID, where every encoder succeeds on 92–100% of episodes. Episode length largely restates success rate, so this panel is informative only at matched success, as here (Sec. IV -G).
College of Computer Science, Sichuan University, Chengdu, China · School of Information and Communication Engineering, University of Electronic Science and Technology of China, Chengdu, China