cs.ROSep 30, 2026

Text-to-3D Policy: Fine-Grained Language-Behavior Alignment for Unseen Specification Generalization

Authors: Xinhao Yang, Wenhao Wu, Ning Lv, Yanshen Ding, Zhenhong Sun, Daoyi Dong, Chunlin Chen, Zhi Wang

Organizations: Nanjing University · Australian National University · University of Technology Sydney

Abstract

3D visuomotor policies provide a strong foundation for spatially precise manipulation, yet current text-to-3D policies struggle to follow unseen fine-grained behavioral specifications beyond those covered by demonstrations. We study this challenge as unseen specification generalization, where language specifies behaviorally significant variations, such as target position, displacement, or articulated state, that are absent from policy training. We find that pretrained language representations and conventional global behavior-language alignment capture coarse task semantics but often blur nearby specifications that require distinct behaviors. We introduce T3DP, a Text-to-3D Policy framework for fine-grained language-behavior alignment. Rather than compressing each instruction and demonstration into a single global embedding, T3DP preserves their local structures and establishes bidirectional token-level correspondence between linguistic elements and behavioral segments. This directly grounds subtle linguistic variations in the behavior components they affect, preventing closely related specifications from collapsing in the representation space. The resulting specification-sensitive language representation conditions a point-cloud-based 3D diffusion policy, enabling more precise control over unseen behavioral specifications without modifying the underlying policy architecture. Across Meta-World, ManiSkill, and RoboTwin, T3DP improves average held-out-specification success over global language-behavior alignment by +11.0-14.2 points, with gains on all 15 task families; on real-robot tasks, it further raises average success from 47.5% to 65.0% (+17.5 points). Representation and action-probe analyses show that fine-grained alignment better preserves specification geometry and action-relevant variation, linking local behavior grounding to downstream control.

Figures & tables

Appendix figures & tables7 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Learning Foresight without Explicit Trajectories for 3D Diffusion Policies

    Sep 17, 2026Zhongbo Zhang, Zaibin Zhang, Yifan Wang +3Diffusion PoliciesAction Generation

  2. Universal Pose Pretraining for Generalizable Vision-Language-Action Policies

    Feb 23, 2026Haitao Lin, Hanyang Yu, Jingshun Huang +5Generalizable Vision-Language-Action Policies

  3. SAM3D-Guided Object-Centric Representation Alignment for Vision-Language-Action Models

    Jul 28, 2026Zonghe Liu, Shanyuan Jie, Xiaoquan Sun +4Diffusion-Based Vision-Language-ActionsRobotic Manipulation