Robot manipulation models primarily reason from 2D observations while acting in the 3D physical world. To bridge this gap, recent work has augmented robot data with geometric priors such as depth, point clouds, and 3D trajectories, while renderable 3D Gaussian representations provide another promising form of 3D supervision. However, 3DGS representation is designed mainly for photometric fidelity and may not preserve real-world metric scale, particularly when the supplied camera extrinsics are unreliable. We study the effect of extrinsic reliability and pose conditioning on feed-forward 3DGS, and propose a calibration-aware pipeline that anchors reconstructed scenes to the robot's metric workspace. Our experiments show that pose conditioning improves novel-view fidelity, while its geometric benefit depends on the reliability of the injected extrinsics. Using this pipeline, we present a renderable, metric-pose-anchored dataset with scene-level reliability information for robot manipulation research. Our dataset is available at https://huggingface.co/datasets/wonguen/3DROID
Figures & tables
Per-scene asset
DROID [ 9 ]
3DROID
Stereo RGB video
✓
✗
Robot joint state
✓
✗
Camera intrinsics †
✓
✓
Shipped extrinsics †
✓
✓
Refined extrinsics ‡
✗
✓
Renderable 3D Gaussians
✗
✓
Table 1: Data availability in DROID and 3DROID. 3DROID adds renderable reconstructions and measured reliability to 114 DROID scenes; source videos and robot states remain in DROID and are joined by episode name. † Intrinsics and shipped extrinsics are from DROID. ‡ PointWorld provides the refined left-eye poses; the corresponding right-eye poses are derived from the known within-rig stereo transform. All camera parameters are packaged in a common convention, with no additional calibration performed by us.
Figure 1: Overview of the 3DROID pose-conditioned construction pipeline. The figure summarizes the scene preparation, calibration gate, reconstruction, and dataset outputs detailed in Section 3 .
Backbone
Method
PSNR ↑
SSIM ↑
LPIPS ↓
YoNoSplat [ 18 ]
Baseline
12.57
0.3582
0.5588
Ours
15.81
0.5105
0.2592
ZipSplat [ 11 ]
Baseline
12.75
0.4233
0.5467
Ours
16.27
0.5242
0.2979
Table 2: Photometric fidelity from three-view reconstructions on 110 scenes. Baseline is pose-free and evaluated with shipped DROID extrinsics. Ours uses refined extrinsics with pose conditioning. Bold indicates the better result within each backbone.
Backbone
Variant
Robot-referenced
Stereo Disagr. (cm) ↓
Proj. Cov.
Err. (cm) ↓
YoNoSplat [ 18 ]
Baseline
0.155
20.38
10.70
Baseline + pose inj.
0.155
18.53
9.93
Baseline + refined ext.
0.167
7.08
7.52
Ours
0.167
6.92
7.54
ZipSplat [ 11 ]
Baseline
0.155
19.71
16.40
Table 3: Geometric consistency of four-view reconstructions. Within each backbone, the upper and lower pairs use shipped ( n=109 ) and refined ( n=110 ) extrinsics, respectively. The extra scene is unmeasurable with shipped extrinsics because all robot queries project off-frame. The second row of each pair injects poses. Because the extrinsic sources define different evaluation geometries, comparisons are made only within each pair.
Figure 2: Qualitative held-out rendering with ZipSplat. Each column shows a different scene. Rows show the ground-truth ext2R image, the pose-free baseline using shipped DROID extrinsics, and our pose-conditioned reconstruction using PointWorld-refined extrinsics. All reconstructions use the remaining three exterior views.
Backbone
Variant
Pose Inj.
Refined Ext.
Photometric Fidelity (3-view)
PSNR ↑
SSIM ↑
LPIPS ↓
AnySplat [ 7 ]
Baseline
✗
✗
8.67
0.3707
0.6641
Baseline + refined ext.
✗
✓
9.60
0.4037
0.5911
YoNoSplat [ 18 ]
Baseline
✗
✗
12.57
0.3582
0.5588
Baseline + pose inj.
✓
✗
15.26
0.4831
0.3074
Baseline + refined ext.
✗
✓
14.45
0.4240
0.3187
Table 4: Ablation of pose conditioning and refined extrinsics on up to 110 scenes. Baseline uses shipped DROID extrinsics without pose conditioning. Baseline + pose injection conditions the backbone on the shipped poses; Baseline + refined extrinsics uses PointWorld-refined extrinsics while keeping the backbone pose-free; Ours combines refined extrinsics with pose conditioning. AnySplat ( n=107 ) is reported only for compatible pose-free configurations because it does not accept external camera poses.
Figure 3: Qualitative ablation on held-out ext2R . Rows show ZipSplat and YoNoSplat on the same scene. After the ground-truth column, reconstruction columns follow the variants defined in Table 4 . All variants use the same three exterior input views.
Ratio-band sweep
Support sweep
Band
Ref.
Ship.
nmin
Ref.
Ship.
±5%
55
40
1
138
88
±10%
99
69
10
127
82
±15%
114
80
30
114
80
±20%
122
89
50
103
68
±30%
122
91
100
91
60
Table 5: Gate-threshold sensitivity over the 225 eligible scenes. The ratio-band sweep fixes nmin=30 , while the support sweep fixes the band to ±15% . Bold denotes the release setting.