3D visuomotor policies provide a strong foundation for spatially precise manipulation, yet current text-to-3D policies struggle to follow unseen fine-grained behavioral specifications beyond those covered by demonstrations. We study this challenge as unseen specification generalization, where language specifies behaviorally significant variations, such as target position, displacement, or articulated state, that are absent from policy training. We find that pretrained language representations and conventional global behavior-language alignment capture coarse task semantics but often blur nearby specifications that require distinct behaviors. We introduce T3DP, a Text-to-3D Policy framework for fine-grained language-behavior alignment. Rather than compressing each instruction and demonstration into a single global embedding, T3DP preserves their local structures and establishes bidirectional token-level correspondence between linguistic elements and behavioral segments. This directly grounds subtle linguistic variations in the behavior components they affect, preventing closely related specifications from collapsing in the representation space. The resulting specification-sensitive language representation conditions a point-cloud-based 3D diffusion policy, enabling more precise control over unseen behavioral specifications without modifying the underlying policy architecture. Across Meta-World, ManiSkill, and RoboTwin, T3DP improves average held-out-specification success over global language-behavior alignment by +11.0-14.2 points, with gains on all 15 task families; on real-robot tasks, it further raises average success from 47.5% to 65.0% (+17.5 points). Representation and action-probe analyses show that fine-grained alignment better preserves specification geometry and action-relevant variation, linking local behavior grounding to downstream control.
Figures & tables
Figure 1: Global alignment matches entire sequences, yielding nearly uniform token-to-segment similarities. Fine-grained alignment learns distinct segment-to-token grounding, enabling more precise control with lower errors.
(a) Target information
Target
Lang.
Succ. ↑
Error ↓
✓
✓
54.7
0.093
✓
✗
55.0
0.091
✗
✓
12.7
0.158
Table 1: Diagnostic study on Meta-World.
Figure 2: Overview of T3DP. (a) A behavior encoder learns specification-dependent representations from demonstration windows through reward prediction conditioned on state–action pairs. (b) With the behavior encoder frozen, the text encoder is trained using global language-behavior alignment and local alignment between state–action hidden states and instruction tokens. (c) The aligned text encoder is frozen to train a language-conditioned 3D diffusion policy using point-cloud observations and proprioceptive states. At inference time, the policy executes unseen specifications from language instructions and current observations, without demonstrations or further adaptation.
Meta-World
Method
Reach
Push
Pick Place
Door Open
Drawer Open
Avg.
DP3 ( Ze et al., 2024 )
9.7 ± 2.3
19.3 ± 4.0
12.7 ± 9.3
15.3 ± 5.0
24.0 ± 4.4
16.2
3DDA ( Ke et al., 2025 )
23.3 ± 18.9
31.7 ± 20.2
26.7 ± 11.5
16.7 ± 2.9
6.7 ± 7.6
21.0
3D-LOTUS ( Garcia et al., 2025 )
21.7 ± 7.6
1.7 ± 2.9
1.7 ± 2.9
38.3 ± 15.3
28.3 ± 2.9
18.3
T2DA ( Zhang et al., 2025 )
24.3 ± 4.0
45.7 ± 8.6
46.0 ± 1.7
41.3 ± 8.3
23.3 ± 2.1
36.1
T3DP (Ours)
38.3 ± 2.9
64.7 ± 17.9
58.0 ± 2.6
59.7 ± 4.2
30.7 ± 12.7
50.3
Table 2: Unseen-specification generalization on Meta-World and ManiSkill. We report the average success rate (%) and corresponding standard deviation over three independently trained seeds. “Avg.” denotes the mean across the five task families within each benchmark. Best and second-best results are shown in bold and underlined , respectively.
Figure 3: Hierarchical component ablation on Meta-World. Each spoke is scaled independently with the task-specific T3DP w/o language score as its inner reference and additional range above the maximum observed score. Avg denotes the mean over the five task families.
Method
# Train Specs.
Reach
Push
Pick Place
Door Open
Drawer Open
Avg.
T2DA
10
22.3
25.3
26.7
28.7
16.7
23.9
20
24.3
45.7
46.0
41.3
23.3
36.1
40
36.1
63.3
75.0
100.0
71.7
69.2
T3DP (Ours)
10
30.3
37.0
34.0
35.3
25.3
32.4
20
38.3
64.7
58.0
59.7
30.7
50.3
40
47.2
75.0
76.7
100.0
83.3
76.4
Table 3: Effect of training specification coverage on unseen-specification generalization in Meta-World. Each training specification contains one expert demonstration, and success rates (%) are evaluated on a common held-out specification set.
Method
Push
Pick Place
Reach
Door Open
Drawer Open
Avg.
DP3
14.3
12.0
1.0
10.3
18.7
11.3
T2DA
42.7
53.0
36.3
31.0
14.3
35.5
T3DP (Ours)
56.7
42.7
35.3
47.3
21.3
40.7
Table 4: Joint multi-task training on Meta-World. A single policy is trained jointly across all five task families and evaluated on held-out specifications within each family. Success rates (%) are averaged over three seeds. Best results are shown in bold .
Figure 5: Real-world manipulation experiments on a PiPer X robotic arm. (a) Language instructions for two task families Pill Bottle and Stapler . (b) Quantitative success rates for DP3, T2DA, and T3DP. (c) Qualitative real-world control results showing the final object positions relative to the marked targets. T3DP achieves a higher success rate and higher final placement precision.
Appendix figures & tables7 assets
Supplementary material from the paper’s appendix.
Appendix
Meta-World
Task
Specification
Instruction Template
Push
Relative displacement
“Push the object [lateral displacement] and [forward displacement].”
Pick Place
Placement target location
“Pick up the object and place it at [target location].”
Reach
End-effector target location
“Move the gripper to [target location].”
Door Open
Target opening extent
“Open the door to [opening extent].”
Drawer Open
Target opening extent
“Open the drawer to [opening extent].”
Appendix
Table 5: Task specifications and language instruction templates across the evaluated benchmarks.
Stage
Configuration
Meta-World
RoboTwin
ManiSkill
I
Training epochs
500
500
500
Batch size
128
128
128
II
Training epochs
375
375
375
Batch size
128
128
128
III
Training epochs
300
600
600
Batch size
512
128
512
Appendix
Table 6: Principal training hyperparameters across the three benchmarks.
Method
Open Microwave
Open Cabinet Drawer
Move Pill Bottle
Move Playing Card
Move Stapler
Avg.
DP3
15.7 ± 8.6
5.0 ± 1.3
3.7 ± 1.3
9.0 ± 2.2
2.7 ± 2.4
7.2
Raw CLIP
37.7 ± 9.5
12.3 ± 10.9
19.3 ± 3.8
28.7 ± 3.4
22.0 ± 2.8
24.0
T2DA
46.0 ± 4.5
16.7 ± 16.3
34.3 ± 1.3
32.7 ± 2.5
24.3 ± 1.7
30.8
T3DP (Ours)
52.0 ± 13.5
37.3 ± 0.6
39.3 ± 0.5
52.3 ± 1.5
28.0 ± 1.6
41.8
Appendix
Table 7: Unseen-specification generalization on RoboTwin. We report success rates (%) and corresponding standard deviations over three independently trained seeds. “Avg.” denotes the mean across the five task families. Best and second-best results are shown in bold and underlined , respectively.
Task
Metric
Generic
Global
Fine-grained
Reach
Spearman ↑
0.504
0.567
0.612
kNN@3 ↑
0.417
0.450
0.533
Push
Spearman ↑
0.255
0.524
0.708
kNN@3 ↑
0.283
0.317
0.567
Door Open
Spearman ↑
0.360
0.682
0.841
kNN@3 ↑
0.400
0.433
0.622
Appendix
Table 8: Representation geometry on held-out Meta-World specifications.
Figure 6: Full same-state action probe results on held-out Push and Reach specifications. Each point compares the predicted and target action under a fixed state while varying the language specification; the dashed diagonal denotes ideal prediction. State-only predictions collapse toward the mean action, whereas behavior–language alignment recovers specification-dependent action variation, with T3DP more closely following the target actions than global alignment (T2DA).
Figure 7: Real-robot setup with a PiPer X manipulator and D435i depth camera.
Figure 8: Spatial distribution of the target specifications used in the real-robot experiments. Left: Move Pill Bottle . Right: Move Stapler . Blue markers denote training specifications, orange markers denote held-out specifications. Each task uses 20 training and 20 held-out target specifications.