3D visuomotor policies provide a strong foundation for spatially precise manipulation, yet current text-to-3D policies struggle to follow unseen fine-grained behavioral specifications beyond those covered by demonstrations. We study this challenge as unseen specification generalization, where language specifies behaviorally significant variations, such as target position, displacement, or articulated state, that are absent from policy training. We find that pretrained language representations and conventional global behavior-language alignment capture coarse task semantics but often blur nearby specifications that require distinct behaviors. We introduce T3DP, a Text-to-3D Policy framework for fine-grained language-behavior alignment. Rather than compressing each instruction and demonstration into a single global embedding, T3DP preserves their local structures and establishes bidirectional token-level correspondence between linguistic elements and behavioral segments. This directly grounds subtle linguistic variations in the behavior components they affect, preventing closely related specifications from collapsing in the representation space. The resulting specification-sensitive language representation conditions a point-cloud-based 3D diffusion policy, enabling more precise control over unseen behavioral specifications without modifying the underlying policy architecture. Across Meta-World, ManiSkill, and RoboTwin, T3DP improves average held-out-specification success over global language-behavior alignment by +11.0-14.2 points, with gains on all 15 task families; on real-robot tasks, it further raises average success from 47.5% to 65.0% (+17.5 points). Representation and action-probe analyses show that fine-grained alignment better preserves specification geometry and action-relevant variation, linking local behavior grounding to downstream control.
Figures & tables
Figure 1: Global alignment matches entire sequences, yielding nearly uniform token-to-segment similarities. Fine-grained alignment learns distinct segment-to-token grounding, enabling more precise control with lower errors.
(a) Target information
Target
Lang.
Succ. ↑
Error ↓
✓
✓
54.7
0.093
✓
✗
55.0
0.091
✗
✓
12.7
0.158
Table 1: Diagnostic study on Meta-World.
Figure 2: Overview of T3DP. (a) A behavior encoder learns specification-dependent representations from demonstration windows through reward prediction conditioned on state–action pairs. (b) With the behavior encoder frozen, the text encoder is trained using global language-behavior alignment and local alignment between state–action hidden states and instruction tokens. (c) The aligned text encoder is frozen to train a language-conditioned 3D diffusion policy using point-cloud observations and proprioceptive states. At inference time, the policy executes unseen specifications from language instructions and current observations, without demonstrations or further adaptation.
Meta-World
Method
Reach
Push
Pick Place
Door Open
Drawer Open
Avg.
DP3 ( Ze et al., 2024 )
9.7 ± 2.3
19.3 ± 4.0
12.7 ± 9.3
15.3 ± 5.0
24.0 ± 4.4
16.2
3DDA ( Ke et al., 2025 )
23.3 ± 18.9
31.7 ± 20.2
26.7 ± 11.5
16.7 ± 2.9
6.7 ± 7.6
21.0
3D-LOTUS ( Garcia et al., 2025 )
21.7 ± 7.6
1.7 ± 2.9
1.7 ± 2.9
38.3 ± 15.3
28.3 ± 2.9
18.3
T2DA ( Zhang et al., 2025 )
24.3 ± 4.0
45.7 ± 8.6
46.0 ± 1.7
41.3 ± 8.3
23.3 ± 2.1
36.1
T3DP (Ours)
38.3 ± 2.9
64.7 ± 17.9
58.0 ± 2.6
59.7 ± 4.2
30.7 ± 12.7
50.3
Table 2: Unseen-specification generalization on Meta-World and ManiSkill. We report the average success rate (%) and corresponding standard deviation over three independently trained seeds. “Avg.” denotes the mean across the five task families within each benchmark. Best and second-best results are shown in bold and underlined , respectively.
Figure 3: Hierarchical component ablation on Meta-World. Each spoke is scaled independently with the task-specific T3DP w/o language score as its inner reference and additional range above the maximum observed score. Avg denotes the mean over the five task families.
Method
# Train Specs.
Reach
Push
Pick Place
Door Open
Drawer Open
Avg.
T2DA
10
22.3
25.3
26.7
28.7
16.7
23.9
20
24.3
45.7
46.0
41.3
23.3
36.1
40
36.1
63.3
75.0
100.0
71.7
69.2
T3DP (Ours)
10
30.3
37.0
34.0
35.3
25.3
32.4
20
38.3
64.7
58.0
59.7
30.7
50.3
40
47.2
75.0
76.7
100.0
83.3
76.4
Table 3: Effect of training specification coverage on unseen-specification generalization in Meta-World. Each training specification contains one expert demonstration, and success rates (%) are evaluated on a common held-out specification set.
Method
Push
Pick Place
Reach
Door Open
Drawer Open
Avg.
DP3
14.3
12.0
1.0
10.3
18.7
11.3
T2DA
42.7
53.0
36.3
31.0
14.3
35.5
T3DP (Ours)
56.7
42.7
35.3
47.3
21.3
40.7
Table 4: Joint multi-task training on Meta-World. A single policy is trained jointly across all five task families and evaluated on held-out specifications within each family. Success rates (%) are averaged over three seeds. Best results are shown in bold .
Figure 5: Real-world manipulation experiments on a PiPer X robotic arm. (a) Language instructions for two task families Pill Bottle and Stapler . (b) Quantitative success rates for DP3, T2DA, and T3DP. (c) Qualitative real-world control results showing the final object positions relative to the marked targets. T3DP achieves a higher success rate and higher final placement precision.
Appendix figures & tables7 assets
Supplementary material from the paper’s appendix.
Appendix
Meta-World
Task
Specification
Instruction Template
Push
Relative displacement
“Push the object [lateral displacement] and [forward displacement].”
Pick Place
Placement target location
“Pick up the object and place it at [target location].”
Reach
End-effector target location
“Move the gripper to [target location].”
Door Open
Target opening extent
“Open the door to [opening extent].”
Drawer Open
Target opening extent
“Open the drawer to [opening extent].”
Appendix
Table 5: Task specifications and language instruction templates across the evaluated benchmarks.
Stage
Configuration
Meta-World
RoboTwin
ManiSkill
I
Training epochs
500
500
500
Batch size
128
128
128
II
Training epochs
375
375
375
Batch size
128
128
128
III
Training epochs
300
600
600
Batch size
512
128
512
Appendix
Table 6: Principal training hyperparameters across the three benchmarks.
Method
Open Microwave
Open Cabinet Drawer
Move Pill Bottle
Move Playing Card
Move Stapler
Avg.
DP3
15.7 ± 8.6
5.0 ± 1.3
3.7 ± 1.3
9.0 ± 2.2
2.7 ± 2.4
7.2
Raw CLIP
37.7 ± 9.5
12.3 ± 10.9
19.3 ± 3.8
28.7 ± 3.4
22.0 ± 2.8
24.0
T2DA
46.0 ± 4.5
16.7 ± 16.3
34.3 ± 1.3
32.7 ± 2.5
24.3 ± 1.7
30.8
T3DP (Ours)
52.0 ± 13.5
37.3 ± 0.6
39.3 ± 0.5
52.3 ± 1.5
28.0 ± 1.6
41.8
Appendix
Table 7: Unseen-specification generalization on RoboTwin. We report success rates (%) and corresponding standard deviations over three independently trained seeds. “Avg.” denotes the mean across the five task families. Best and second-best results are shown in bold and underlined , respectively.
Task
Metric
Generic
Global
Fine-grained
Reach
Spearman ↑
0.504
0.567
0.612
kNN@3 ↑
0.417
0.450
0.533
Push
Spearman ↑
0.255
0.524
0.708
kNN@3 ↑
0.283
0.317
0.567
Door Open
Spearman ↑
0.360
0.682
0.841
kNN@3 ↑
0.400
0.433
0.622
Appendix
Table 8: Representation geometry on held-out Meta-World specifications.
Figure 6: Full same-state action probe results on held-out Push and Reach specifications. Each point compares the predicted and target action under a fixed state while varying the language specification; the dashed diagonal denotes ideal prediction. State-only predictions collapse toward the mean action, whereas behavior–language alignment recovers specification-dependent action variation, with T3DP more closely following the target actions than global alignment (T2DA).
Figure 7: Real-robot setup with a PiPer X manipulator and D435i depth camera.
Figure 8: Spatial distribution of the target specifications used in the real-robot experiments. Left: Move Pill Bottle . Right: Move Stapler . Blue markers denote training specifications, orange markers denote held-out specifications. Each task uses 20 training and 20 held-out target specifications.
3D diffusion policies are strong at generating geometrically grounded actions from current observations, but successful manipulation requires not only knowing what motion is feasible now, but also anticipating where the interaction is heading. Existing policies largely leave such foresight to emerge implicitly from action learning. We introduce Movement Trend Guidance, a simple but effective way to provide this foresight without introducing an explicit plan. From a short observation history, the policy learns a compact latent representation of interaction evolution. During training, sparse future gripper states supervise this representation; at inference, only the latent is retained as future-oriented conditioning alongside the current observation. The latent provides global conditioning for action generation, while an additional gated FiLM branch is used only at the UNet bottleneck. Despite adding only 3.52% more parameters to DP3, our method preserves the original dense-action and receding-horizon formulation and consistently improves upon DP3 across RoboTwin2.0, LIBERO-40, and DexArt. It reaches 62.8% vs. 56.1% in 50-task RoboTwin2.0 mixed training, 71.93% vs. 37.08% on LIBERO-40, and 72.0% vs. 49.0% on five real-robot tasks. These results show that a diffusion policy can benefit substantially from knowing where an interaction is heading, without being told exactly where to move.
Existing Vision-Language-Action (VLA) models often suffer from feature collapse and low training efficiency because they entangle high-level perception with sparse, embodiment-specific action supervision. Since these models typically rely on VLM backbones optimized for Visual Question Answering (VQA), they excel at semantic identification but often overlook subtle 3D state variations that dictate distinct action patterns. To resolve these misalignments, we propose Pose-VLA, a decoupled paradigm that separates VLA training into a pre-training phase for extracting universal 3D spatial priors in a unified camera-centric space, and a post-training phase for efficient embodiment alignment within robot-specific action space. By introducing discrete pose tokens as a universal representation, Pose-VLA seamlessly integrates spatial grounding from diverse 3D datasets with geometry-level trajectories from robotic demonstrations. Our framework follows a two-stage pre-training pipeline, establishing fundamental spatial grounding via poses followed by motion alignment through trajectory supervision. Extensive evaluations demonstrate that Pose-VLA achieves state-of-the-art results on RoboTwin 2.0 with a 79.5% average success rate and competitive performance on LIBERO at 96.0%. Real-world experiments further showcase robust generalization across diverse objects using only 100 demonstrations per task, validating the efficiency of our pre-training paradigm.
Haitao Lin, Hanyang Yu, Jingshun Huang +5
Tencent Robotics X · Futian Laboratory · The Hong Kong University of Science and Technology +2
Vision-Language-Action (VLA) models have shown strong potential for general robot manipulation, but most existing models rely on 2D visual-language backbones and lack fine-grained 3D understanding of target objects, especially under occlusion, pose variation, scale changes, and precise spatial interaction. We propose an object-centric 3D representation alignment framework built upon π0, using SAM3D as a frozen 3D teacher to provide target-object 3D priors during training. Specifically, we localize task-relevant objects with object recognition models, generate corresponding object masks, and use SAM3D to extract dense object-level 3D representations, which are aligned with intermediate visual features of π0. This enables the policy to internalize target-object 3D information while preserving the original RGB-language-to-action inference pipeline without requiring depth, point clouds, masks, SAM3D, or additional 3D modules at test time. Simulation experiments show consistent improvements, achieving 99.1% on LIBERO and an average length of 4.11 on CALVIN. Real-world experiments further demonstrate that our method is particularly effective in long-horizon manipulation scenarios where the robot must focus on different target objects across multiple subtasks.
Zonghe Liu, Shanyuan Jie, Xiaoquan Sun +4
University of Hong Kong · Shenzhen Institutes of Advanced Technology, Chinese Academy of Sciences · Huazhong University of Science and Technology +2