Scaling up robotic manipulation is primarily bottlenecked by the scarcity of real-world robot data. While recent approaches leverage human video demonstrations to mitigate this shortage, they remain computationally expensive and still rely on paired human-robot data for domain alignment. Although current state-of-the-arts excel at long-horizon tasks, they struggle with the delicate and precise control required for complex tool manipulation. To overcome these limitations, we introduce P2P-T, from Pixel to Poses for Tool Manipulation, a data-efficient, object-centric framework that learns tool use directly from human demonstrations. P2P-T bridges the cognitive and physical execution gap through a two-stage approach. First, pretraining an object-centric world model to extract stable pose priors; second, integrating these priors into an efficient, pose-aware low-level policy. By utilizing a robust automated data processing pipeline powered by modern foundation models, P2P-T completely bypasses the need for human-robot aligned data. This reduces overall training overhead drastically. With minimal per-task fine-tuning, our framework achieves a 73% improvement over the previous state of the art in execution performance on complex, real-world tool manipulation tasks that currently remain out of reach for standard large-scale pretrained models.
Figures & tables
Fig. 3: Overview of the P2P-T pre-processing pipeline. We first extract the tool from the background environment to create a 2D segmentation mask. The mask and the raw RGB frames are subsequently fed into SAM 3D [ 2 ] , which outputs a high-fidelity, instance-specific 3D tool mesh. Finally, FoundationPose [ 3 ] takes the reconstructed 3D mesh and the sequence of raw video frames as inputs to compute and track precise 6D tool poses across the entire demonstration sequence.
Fig. 4: Overview of P2P-T’s two-stage framework. Stage 1: An autoregressive transformer learns an object-centric world model from human demonstrations by fusing observations encoded with SigLIP [ 20 ] and DINOv2 [ 21 ] , instructions encoded with T5 [ 22 ] , and projected tool poses. It learns future pose latents spanning t+1 through t+H−1 , supervised through decoded poses and auxiliary frame predictions. Stage 2: The frozen world model provides future pose latents, mapped through a trainable adaptor to condition an RDT-based diffusion action expert. Together with visual observations, language instructions, and robot proprioception, these latents guide action chunks at:t+H−1 for receding-horizon execution, connecting human-derived object-centric dynamics with robot-specific control without paired human–robot demonstrations.
Fig. 5: Real-world evaluation tasks. We evaluate on six complex tool-use scenarios that demand delicate physical control and continuous spatial reasoning, moving far beyond standard pick-and-place actions.
Method
Hammer
Cup
Brush
Screwdriver
Knife
Spoon
Average ( ↑ )
Execution Rate (Hz) ( ↑ )
Vision-Language-Action Models
DP2-DINOv2 [ 25 ]
0.09
0.12
0.03
0.02
0.09
0.04
0.07
30
Octo [ 10 ]
0.21
0.27
0.27
0.17
0.16
0.26
0.22
15
OpenVLA-OFT [ 26 ]
0.09
0.27
0.13
0.17
0.24
0.12
0.17
25
π0.5 [ 13 ]
0.24
0.49
0.27
0.35
0.37
0.24
0.33
30
World Action Models
TABLE I: Real-world Policy Evaluation. We report the success rate for each task, together with the average success rate and inference rate. Each model is fine-tuned with 100 demonstrations and evaluated over 50 trials per task. We report the average results, with the best success rates highlighted in bold . Inference rate is measured in Hz, with higher values being better.
Object-centric Pretraining
Pose-aware Post-training
Hammer
Cup
Average
✓
✗
0.12
0.14
0.13
✗
✓
0.32
0.36
0.34
✓
✓
0.59
0.60
0.60
TABLE II: Ablations on Framework Design.
Model Variant
#Param.
Hammer
Cup
Average
P2P-T-Small
1.189B
0.34
0.48
0.41
P2P-T-Base
1.921B
0.59
0.60
0.60
P2P-T-Large
2.832B
0.64
0.68
0.66
TABLE III: Scaling Verification. Performance comparison among the P2P-T-Small, P2P-T-Base, and P2P-T-Large variants. “#Param.” denotes the number of parameters.