Visuomotor policies for humanoid loco-manipulation must generalize across object poses and lighting from only a handful of demonstrations. 3D Diffusion Policy (DP3) conditions a diffusion-based action generator on point-cloud features, but its PointNet-style encoder has no built-in equivariance to the rotations, translations, and scalings (SIM(3)) that manipulation tasks respect. EquiBot closed this gap for wheeled manipulators with a SIM(3)-equivariant Vector Neuron Network (VNN) encoder. We extend this to a substantially more complex embodiment, the 43-joint Unitree G1 humanoid, and propose EquivDP3: a two-stage policy where a high-level diffusion planner with a SIM(3)-equivariant VNN encoder emits 6 Hz whole-body command chunks, executed at 50 Hz by a frozen, pre-trained RL locomotion policy and a differential inverse-kinematics module for the arms, trained end-to-end by behavior cloning. Across two simulated IsaacLab benchmarks and four non-equivariant baselines (5-100 demonstrations, in- and out-of-distribution), EquivDP3's advantage concentrates in the low-data regime: at 5-10 demonstrations it reaches 67.1% success versus 38.2-52.4% for the baselines, while by 50-100 all encoders converge (74.3-85.2%) and the ordering is no longer meaningful. A proprioception-only control confirms this gap is genuinely perceptual: with the point cloud removed, success drops to 31% vs. 60% (EquivDP3) at 5 demonstrations and 78% vs. 99% at 10, but vanishes by 50-100, showing the high-data plateau reflects a benchmark ceiling, not five encoders learning the same invariance. The encoder costs only 0.8 ms of extra latency per action chunk over the PointNet encoder it replaces. Baking geometric symmetry into a hierarchical diffusion policy's perception backbone is a practical, nearly free way to improve data efficiency for humanoid loco-manipulation when demonstrations are scarce.
Figures & tables
Fig. 1 : EquivDP3 two-stage architecture. A diffusion planner with a SIM(3)-invariant point-cloud encoder runs at ∼ 6 Hz and emits whole-body command chunks. A low-level controller at 50 Hz splits each command between an RL-trained locomotion policy (velocity/angular/height/torso → leg joints) and an inverse-kinematics module (wrist pose → arm joints); both feed a joint-space PD controller.
Encoder
Rotation
Translation
Scale
EquivDP3 (ours)
3.9×10−7
2.7×10−7
3.8×10−7
PointNet
2.3×10−1
5.3×10−1
1.5×10−1
Ours, no LayerNorm
4.1×10−7
6.0×10−7
5.2
TABLE I : Empirical invariance check. Relative error ∥f(g⋅P)−f(P)∥/∥f(P)∥ over random transforms; 10−7 is float32 round-off. The final row shows that disabling the encoder’s LayerNorms removes scale invariance only.
Fig. 2 : Overview of the two evaluation tasks. Task 1 (Galileo Loco-Manipulation, left) requires full whole-body coordination across four subgoals: approach/walk to the shelf, bimanual pick, rotate & carry, and place into the sorting bin. Task 2 (G1 Pick-and-Place, right) fixes the base and requires only two subgoals: bimanual pick and place into the target bin.
Task
Subgoals
# subgoals
Galileo Loco-Manip.
approach/walk to shelf → bimanual pick → rotate & carry → place into bin
4
G1 Pick-and-Place
pick → place (fixed base)
2
TABLE II : Evaluation tasks.
Fig. 3 : Overall success rate vs. number of demonstrations, for Task 1 (Galileo Loco-Manipulation, left column) and Task 2 (G1 Pick-and-Place, right column), under in-distribution (top row) and out-of-distribution (bottom row) conditions. EquivDP3 (teal) leads in aggregate at 5–10 demonstrations; by 50–100 all encoders converge, and it is not the top performer in every individual setting.
EquivDP3
PointNet
PointNet+Aug
DGCNN
PCT
Task
ndemo
ID
OOD
ID
OOD
ID
OOD
ID
OOD
ID
OOD
Task 1
5
79.00 ± 7.91
47.00 ± 9.60
43.00 ± 9.53
49.00 ± 9.61
73.00 ± 8.58
40.00 ± 9.43
17.00 ± 7.33
20.00 ± 7.77
4.00 ± 4.14
2.00 ± 3.23
10
89.00 ± 6.19
60.00 ± 9.43
53.00 ± 9.60
37.50 ± 6.65
81.00 ± 7.63
59.00 ± 9.47
4.00 ± 4.14
12.00 ± 6.41
89.00 ± 6.19
66.00 ± 9.13
50
97.00 ± 3.71
65.00 ± 9.19
94.00 ± 4.85
61.50 ± 6.68
98.00 ± 3.23
78.00 ± 8.03
94.00 ± 4.85
64.00 ± 9.25
75.00 ± 5.96
46.00 ± 9.59
100
98.00 ± 3.23
86.00 ± 6.81
97.00 ± 3.71
70.00 ± 6.30
99.00 ± 2.64
84.00 ± 7.16
100.00 ± 1.85
71.00 ± 8.76
100.00 ± 1.85
78.00 ± 8.03
Task 2
5
60.00 ± 9.43
27.00 ± 8.58
10.00 ± 4.19
17.00 ± 7.33
34.00 ± 6.51
21.00 ± 7.91
85.00 ± 6.99
66.00 ± 9.13
42.33 ± 5.56
16.67 ± 5.95
TABLE III : Success rate (%) by encoder, task, and demonstration count, with in-distribution (ID) and out-of-distribution (OOD) as paired sub-columns under each encoder. Each cell reports the point estimate and the half-width of its 95% Wilson score interval [ 33 ] ( p±h ), computed directly from the logged per-episode outcomes. Episodes per cell: 63 cells use 100, 6 use 150, 9 use 200, 2 use 300. Cells evaluated more than once on the identical checkpoint are pooled, with n the pooled count; which cells were repeated reflects development history, not selection on outcome.
ndemo=5
ndemo=10
Encoder
SG1
SG2
SG3
SG4
SG1
SG2
SG3
SG4
EquivDP3
100
90
87
79
100
98
97
89
PointNet
100
32
27
43
100
79
53
53
PointNet+Aug
100
77
76
73
100
99
98
81
DGCNN
100
42
29
17
100
100
100
4
PCT
100
18
17
4
100
99
96
89
TABLE IV : Task 1 subgoal event rates (%), ID, at the two low-data budgets. SG1 approach, SG2 lift, SG3 carry to bin (gated on SG2), SG4 task success (not gated on SG2).
Fig. 4 : ID → OOD success-rate drop (percentage points) for both tasks, at (a) ndemo=100 and (b) ndemo=10 . EquivDP3 shows the smallest drop among all encoders at n=100 on Task 1, and DGCNN nearly collapses on Task 1 at n=10 (both ID and OOD performance near floor).
n=5
n=10
n=50
n=100
In-distribution
Blind (no point cloud) a
31
78
92
94
EquivDP3 (ours)
60
99
90
96
Gap (EquivDP3 − Blind)
+29
+21
−2
+2
Out-of-distribution
Blind (no point cloud) a
21
62
65
64
TABLE V : Proprioception-only control on Task 2 across the full demonstration-budget grid of Table III , under both distribution conditions. Bold gaps are significant by Fisher’s exact test ( p<0.05 ).
Fig. 5 : Left: median per-chunk inference latency over all single-environment evaluation runs, well under the 160 ms real-time budget (dashed line) implied by an 8-step chunk at 50 Hz; EquivDP3 costs 0.8 ms more than PointNet, and PCT is the slowest. Right: mean episode length ( ± std) in simulated steps at ndemo=100 , ID, where every encoder succeeds on 92–100% of episodes. Episode length largely restates success rate, so this panel is informative only at matched success, as here (Sec. IV -G).
Diffusion-based visuomotor policies operating directly in raw action spaces conflate scene comprehension with trajectory generation within a single denoising process. The resulting velocity field must simultaneously encode scene information and generate precise trajectories, increasing learning complexity and limiting performance on tasks demanding precise temporal coordination across multiple arms. To simplify this joint learning problem, we introduce Latent Diffusion Policy (LDP), a two-stage framework performing flow matching in a deliberately shaped latent space. By absorbing scene understanding into an observation-conditioned CVAE encoder, LDP concentrates the conditional distribution of each observation. Consequently, the flow model avoids implicitly resolving scene-dependent structures; instead, it generates within a pre-concentrated distribution featuring a smoother velocity field, simplifying learning from limited demonstrations. Furthermore, to capture temporal dependencies among latent tokens, LDP trains with per-token diffusion forcing and employs staircase inference sampling to resolve the resulting distributional mismatch. We also propose reconstruction FID (rFID) as a lightweight proxy predicting downstream task success solely from latent space statistics. On coordination-intensive tasks from RoboTwin 2.0, LDP outperforms DP3 by a substantial margin and transfers effectively to real-world bimanual deployments.
Diffusion-based visuomotor policies model complex action distributions through iterative denoising, but repeated inference adds latency to robotic control. One-step generators reduce this cost, motivating training objectives that retain useful action structure with few demonstrations. We present Ada3Drift, a point-cloud-conditioned policy that builds on Drifting Models to perform distribution refinement during training and generate action chunks in one forward pass. Our central design is a regression-to-drifting curriculum: paired action regression first emphasizes observation--action correspondence, while a sigmoid schedule progressively increases a batch-level action-distribution regularizer. The drifting term combines attraction to demonstrated actions and repulsion among generated samples using inherited multi-temperature aggregation. A timestep-free generator preserves single-step inference throughout. On Adroit, Meta-World, RoboTwin, and five real-world tasks, Ada3Drift achieves the highest reported average success rates among the evaluated baselines with 1 NFE, compared with 10 NFE for the diffusion baselines. Controlled ablations favor sigmoid scheduling at matched demonstration budgets. We will release our code and pretrained model weights.
Chongyang Xu, Yixian Zou, Tianyu Yang +4
College of Computer Science, Sichuan University, Chengdu, China · School of Information and Communication Engineering, University of Electronic Science and Technology of China, Chengdu, China
Diffusion-based visuomotor policies perform well in robotic manipulation, yet current methods still inherit image-generation-style decoders and multi-step sampling. We revisit this design from a frequency-domain perspective. Robot action trajectories are highly smooth, with most energy concentrated in a few low-frequency discrete cosine transform modes. Under this structure, we show that the error of the optimal denoiser is bounded by the low-frequency subspace dimension and residual high-frequency energy, implying that denoising error saturates after very few reverse steps. This also suggests that action denoising requires a much simpler denoising model than image generation. Motivated by this insight, we propose Hyper-DP3 (HDP3), a pocket-scale 3D diffusion policy with a lightweight Diffusion Mixer decoder that supports two-step DDIM inference. Our synthetic experiments validate the theory and support the sufficiency of two-step denoising. Futhermore, across RoboTwin2.0, Adroit, MetaWorld, and real-world tasks, HDP3 achieves state-of-the-art performance with fewer than 1% of the parameters of prior 3D diffusion-based policies and substantially lower inference latency.
Jinhao Zhang, Zhexuan Zhou, Huizhe Li +5
Harbin Institute of Technology, Shenzhen · Shanghai Jiao Tong University