DeltaSplat: Iterative Gaussian Refinement for Pose-Free Feed-Forward 3D Gaussian Splatting
Authors: Chanung Park, Seunghyeon Song, Joo Chan Lee, Eunbyung Park, Jong Hwan Ko
Organizations: Department of Electrical and Computer Engineering, Sungkyunkwan University · Department of Artificial Intelligence, Sungkyunkwan University · Department of Artificial Intelligence, Yonsei University
Pose-free feed-forward 3D Gaussian Splatting (3DGS) reconstructs a scene from sparse, unposed images in a single network pass, removing the need for camera calibration and per-scene optimization. However, camera estimation errors propagate into the predicted Gaussians and compound the geometric and photometric inaccuracies of single-pass prediction. To correct these errors, we introduce DeltaSplat, a lightweight Gaussian refinement module for pose-free feed-forward 3DGS. It iteratively renders the current Gaussians at the input context views and predicts per-Gaussian updates from the resulting residuals. A 2D residual alone, however, underdetermines the 3D correction. DeltaSplat therefore conditions each update on per-pixel Plücker rays and rendered depth as a soft geometric prior. A dual-branch convolutional mixer efficiently encodes these inputs, and per-attribute heads decode the fused features into position, opacity, and color updates. The module adds only ~2.2% parameters to the backbone and remains fully feed-forward at inference. On DL3DV, DeltaSplat reaches 26.64 dB PSNR in the pose-free setting, improving its state-of-the-art backbone by 1.75 dB and surpassing even baselines supplied with ground-truth cameras; consistent gains hold across 6-24 views and all camera regimes.
Figures & tables
Figure 2: Method overview. The backbone predicts cameras and initial Gaussians; at each iteration, the refinement module encodes the render residual with the Plücker rays and the rendered depth, and decodes the per-Gaussian update {Δμ,Δc,Δα} , repeated K=3 times with shared weights.
Figure 3: Geometry-guided residual encoding. (a) The 2D residual alone underdetermines the position update Δμ : displacements along the viewing ray yield identical residuals, leaving a large possible region (red). (b) The per-pixel Plücker ray and rendered depth encode the line of sight and the current prediction’s position along it, a prior on the current geometry under which the possible region contracts.
6 views
12 views
24 views
Method
p
K
PSNR ↑
SSIM ↑
LPIPS ↓
PSNR ↑
SSIM ↑
LPIPS ↓
PSNR ↑
SSIM ↑
LPIPS ↓
MVSplat ( Chen et al., 2024 )
✓
✓
22.66
0.760
0.173
21.29
0.709
0.224
19.98
0.662
0.269
DepthSplat ( Xu et al., 2025b )
✓
✓
23.42
0.797
0.136
21.91
0.753
0.179
20.09
0.690
0.240
YoNoSplat ( Ye et al., 2026 )
✓
✓
24.72
0.817
0.139
23.29
0.773
0.177
22.67
0.758
0.192
DeltaSplat
✓
✓
26.70
0.866
0.113
25.33
0.833
0.142
24.72
0.824
0.151
NoPoSplat ( Ye et al., 2025 )
✓
22.77
0.743
0.179
19.38
0.563
0.318
17.86
0.495
0.397
Table 1: Main results on the DL3DV testset (140 scenes, 2242 ) at 6 / 12 / 24 context views. p , K : ✓ means the ground-truth pose / intrinsics are supplied, blank means predicted by the model.
Figure 4: Qualitative results on DL3DV scenes.
32 views
64 views
128 views
Method
PSNR ↑
SSIM ↑
LPIPS ↓
PSNR ↑
SSIM ↑
LPIPS ↓
PSNR ↑
SSIM ↑
LPIPS ↓
YoNoSplat
17.94
0.659
0.380
18.83
0.688
0.342
19.28
0.701
0.325
DeltaSplat
18.62
0.670
0.371
19.76
0.711
0.328
20.62
0.741
0.298
Table 2: Cross-dataset generalization to ScanNet++ ( Yeshwanth et al., 2023 ) at 2242 with 32 / 64 / 128 context views and known intrinsics. All models are trained on DL3DV.
Figure 6
Figure 6: Effect of the 3D geometry prior. Each pair compares the geometry-aimed encoding with an encoding built from the 2D residual alone; the left pair shows the refined render and the right pair the per-Gaussian ∣Δμ∣ . Without the Plücker-ray and depth channels the same residual drives larger and more scattered position updates, and PSNR falls from 25.1 to 20.3 dB.
Table 4: Ablations on the DL3DV test set in the pose-free setting with known intrinsics (6 context views). (a) Each row removes a single component from the full model. (b) Effect of the number of refinement iterations K applied at test time, with per-scene refinement time; K=0 denotes the unrefined backbone output, and the model is trained with K=3 .
Feed-forward 3D Gaussian Splatting methods reconstruct a scene from posed or pose-free images in a single forward pass, yet current approaches predict one Gaussian per input pixel, tying the representation budget to camera resolution rather than scene complexity. A flat wall and a richly textured object thus produce equally many Gaussians despite very different geometric needs. We propose ZipSplat, a token-based feed-forward model that decouples Gaussian placement from the pixel grid. A multi-view backbone extracts dense visual tokens, and k-means clustering compresses them into a compact set of scene tokens. Cross- and self-attention refine these tokens, and a lightweight MLP decodes each into a group of Gaussians with unconstrained 3D positions. Because clustering is applied at inference, a single trained model spans the quality-efficiency curve without retraining. ZipSplat operates without ground-truth poses or intrinsics, yet sets a new state of the art on DL3DV and RealEstate10K with ∼6× fewer Gaussians than pixel-aligned methods, surpassing the best pose-free baseline by 2.1dB and 1.2dB PSNR, respectively. It further generalizes zero-shot to Mip-NeRF360 and ScanNet++, outperforming all comparable baselines. Our project page is at https://veichta.com/zipsplat.
We present StructSplat, a feed-forward and generalizable 3D Gaussian reconstruction framework that operates directly on uncalibrated images without requiring camera parameters. Existing methods either rely on per-scene optimization or assume known camera poses, and often entangle geometry and appearance within a unified backbone, limiting reconstruction fidelity and generalization. Our key idea is to adopt a structured representation that organizes geometry, semantic, and texture cues with explicit roles in the reconstruction process. Specifically, we introduce a pixel-aligned feature injection mechanism to enable accurate texture modeling from 2D observations, incorporate semantic-aware priors to improve global consistency, and design a camera alignment strategy to prevent information leakage and improve generalization. Experiments show that our method significantly outperforms prior approaches on challenging benchmarks. On DL3DV, our method achieves 28.045 PSNR, surpassing AnySplat (22.377) by +5.67 dB. In cross-dataset evaluation, our method achieves +1.94 dB over AnySplat on ACID and +1.72 dB on RealEstate10K. Project page: https://structsplat.github.io Code: https://github.com/J-C-Zhao/StructSplat
Jia-Chen Zhao, Beiqi Chen, Xinyang Chen +2
Harbin Institute of Technology (Shenzhen) · Great Bay University · Guangzhou CloudButterfly Technology Co., Ltd.
Feed-forward 3D Gaussian Splatting models offer fast single-pass reconstruction,but scaling them to match per-scene optimization quality is fundamentally hindered by the scarcity of large-scale 3D annotations. A practical compromise is predict-then-refine,where post-prediction optimization compensates for the limited capacity of the feed-forward network. However,standard feed-forward 3DGS is trained solely for zero-step rendering error,ignoring whether its output constitutes a good initialization for the downstream optimizer. We present ForeSplat,an optimization-aware training framework that equips feed-forward 3DGS models to produce initializations explicitly designed for rapid,effective refinement. By offloading part of the scene-modeling burden to the optimizer,ForeSplat substantially reduces the capacity pressure on the feed-forward model,making high-quality reconstruction feasible even with compact networks. At its core is MetaGrad,a lightweight multi-anchor meta-gradient training rule that bypasses costly higher-order differentiation through the 3DGS optimizer. MetaGrad unrolls a short inner-loop refinement trajectory,samples anchor states,and back-propagates aggregated first-order gradients to the prediction head as a surrogate optimization-aware signal. This fine-tuning adds no inference cost and enables high-quality reconstruction within seconds after a few refinement steps. We instantiate ForeSplat on diverse backbones,including AnySplat,Pi3X,and a distilled variant tailored for edge deployment. Across all tested architectures,a ForeSplat-trained initialization converges in fewer refinement steps and reaches a higher peak reconstruction quality than its vanilla counterpart,even fully converged. The framework consistently bridges the gap between amortized prediction and per-scene optimization,establishing a practical path toward lightweight,high-fidelity 3D reconstruction.
Yuke Li, Weihang Liu, Cheng Zhang +8
ShanghaiTech University · GGU Technology Co., Ltd · Stereye