Pose-free feed-forward 3D Gaussian Splatting (3DGS) has demonstrated remarkable potential for generalized novel view synthesis. However, existing methods typically predict Gaussian appearance attributes represented by spherical harmonics (SH) in the same manner, overlooking the fundamental distinction between view-independent and view-dependent appearance, which results in suboptimal rendering quality. In this paper, we present AESplat, a novel and general framework for pose-free feed-forward 3DGS that introduces an effective decoupled appearance modeling strategy based on an analysis of SH, enabling higher-quality rendering. Specifically, AESplat directly derives the zeroth-order SH coefficient, which represents the base view-independent appearance component, from the input images without training. The higher-order SH coefficients are subsequently predicted by a shallow multilayer perceptron equipped with two efficient 3D-aware inductive biases to model view-dependent appearance variations. Extensive experiments across multiple datasets demonstrate that our method significantly outperforms state-of-the-art approaches, achieving a 0.8 dB improvement in PSNR over the pose-free method NAS3R and a 1.1 dB improvement over the pose-required method DepthSplat on the RealEstate10K dataset. Project page: https://aesplat.github.io/.
Figures & tables
Figure 1: Our method AESplat consistently outperforms state-of-the-art pose-free feed-forward methods in rendering quality across both indoor and outdoor scenes. Yellow and red boxes highlight differences in diffuse (e.g., walls) and specular (e.g., mirrors and glass) appearance, respectively.
Figure 2: Architectural Differences. Unlike (a) existing pose-free methods that predict all appearance attributes with a unified prediction head, (b) we adopt a decoupled formulation that employs tailored strategies to model view-independent and view-dependent appearance separately.
Figure 3: Pipeline of AESplat . (a) Given unposed images, AESplat jointly predicts camera poses and Gaussian primitives in a single forward pass, with their appearance attributes obtained through decoupled appearance modeling. (b) Leveraging the strong observation provided by the input image itself, the zeroth-order SH coefficient of each pixel-aligned Gaussian is directly derived from its corresponding pixel RGB value. (c) Higher-order SH coefficients are predicted by a view-dependent appearance head, which takes two effective 3D-aware inductive biases, GVSRE and WCM, enabling more accurate modeling of view-dependent appearance.
Method
RealEstate10K
ACID
PSNR ↑
SSIM ↑
LPIPS ↓
PSNR ↑
SSIM ↑
LPIPS ↓
Supervised Pose-required
pixelSplat ( Charatan et al., 2024 )
23.859
0.808
0.184
25.889
0.780
0.194
MVSplat ( Chen et al., 2024a )
24.012
0.812
0.175
25.561
0.775
0.195
DepthSplat ( Xu et al., 2025b )
25.595
0.852
0.145
−
−
−
YoNoSplat ( Ye et al., 2026 )
24.233
0.813
0.162
−
−
−
Table 1: Performance comparison of novel view synthesis on RealEstate10K and ACID datasets . We report the average metrics across all test scenes. The best and second-best results are highlighted. − indicates that the result was not reported in the original paper.
Figure 4: Qualitative results of Tab. 1 . The leftmost column shows the two-view context images. The top two rows are from RealEstate10K, and the bottom two are from ACID.
Method
ACID
DL3DV
ScanNet++
PSNR ↑
SSIM ↑
LPIPS ↓
PSNR ↑
SSIM ↑
LPIPS ↓
PSNR ↑
SSIM ↑
LPIPS ↓
Supervised Pose-required
pixelSplat
25.477
0.770
0.207
18.688
0.582
0.354
18.422
0.720
0.278
MVSplat
25.525
0.773
0.199
17.786
0.545
0.357
17.138
0.687
0.297
DepthSplat
26.012
0.791
0.185
19.553
0.611
0.285
20.775
0.760
0.254
YoNoSplat
24.246
0.721
0.222
19.636
0.594
0.311
21.075
0.744
0.254
Table 2: Cross-dataset generalization . All methods are trained on RealEstate10K and evaluated in a zero-shot setting on ACID, DL3DV, and ScanNet++.
Figure 5: Qualitative results of Tab. 2 . From top to bottom, the three rows show zero-shot results on ACID, DL3DV, and ScanNet++ when trained on RealEstate10K.
Figure 8
Figure 7: Qualitative results of Tab. 4 . GVSRE and WCM model view-dependent appearance variations (red boxes), while I2DC better preserves the base appearance component (yellow boxes).
Method
PSNR ↑
SSIM ↑
LPIPS ↓
Full
24.268
0.821
0.160
(1) w/o GVSRE
23.999
0.816
0.163
(2) w/o WCM
24.012
0.818
0.161
(3) I2DC only
23.610
0.808
0.173
(4) Baseline (NAS3R)
23.754
0.808
0.171
(5) Baseline (NAS3R) + I2DC
23.908
0.813
0.165
Table 4: Ablations. We evaluate the contribution of the proposed method on RealEstate10K.
While feed-forward 3D Gaussian Splatting (3DGS) enables efficient 3D reconstruction, achieving high-fidelity rendering remains challenging. Existing pixel-aligned approaches suffer from spatial inflexibility and massive structural redundancy, whereas query-based methods lack 3D priors and entangle geometry with appearance, yielding blurry, pose-dependent results. To overcome these deficiencies, we propose \textbf{QuerySplat}, a feed-forward 3DGS framework driven by geometric priors and explicit appearance decoupling. Specifically, we design a dual-branch query-based decoder: the geometry branch leverages a pretrained Vision Geometric Model for spatial understanding, which intrinsically endows QuerySplat with pose-free modeling capabilities, while the appearance branch recovers high-frequency details through a dedicated pathway separated from geometric attribute regression. Extensive experiments demonstrate that QuerySplat mitigates the blurry rendering issues of early query-based models and consistently outperforms pixel-aligned approaches in rendering fidelity. On the challenging DL3DV benchmark, it achieves state-of-the-art novel view synthesis performance, with average PSNR gains of 2.30 dB and 1.04 dB over the best pose-free and pose-required baselines, respectively. Project Page: https://inspatio.github.io/querysplat.
Yinglong Li, Donghui Shen, Xiaoyu Zhang +5
State Key Laboratory of Virtual Reality Technology and Systems, Beihang University · InSpatio Research · State Key Lab of CAD&CG, Zhejiang University
We present ViewSplat, a view-adaptive 3D Gaussian splatting network for novel view synthesis from unposed images. While recent feed-forward 3D Gaussian splatting has significantly accelerated 3D scene reconstruction by bypassing per-scene optimization, a fundamental fidelity gap remains. We attribute this gap to the limited capacity of single-step feed-forward networks to regress static Gaussian primitives that satisfy all viewpoints. To address this limitation, we shift the paradigm from static primitive regression to view-adaptive splatting. Instead of a rigid Gaussian representation, our pipeline learns a view-adaptive latent representation. Specifically, ViewSplat initially predicts base Gaussian primitives alongside the weights of scene-conditioned View MLPs. During rendering, these MLPs take target-view coordinates as input and predict view-dependent residual updates for each Gaussian attribute (i.e., 3D position, scale, rotation, opacity, and color). This mechanism, which we term view-adaptive splatting, allows each primitive to rectify initial estimation errors, effectively capturing high-fidelity appearances. Extensive experiments demonstrate that ViewSplat achieves state-of-the-art fidelity while maintaining fast inference and real-time rendering; our large backbone variant runs at 15 FPS during inference and 90 FPS during rendering. Our project page is available at https://cvlab-uos.github.io/ViewSplat.
Moonyeon Jeong, Seunggi Min, Suhyeon Lee +1
University of Seoul · Korea Electronics Technology Institute
3D Gaussian Splatting (3DGS) achieves high-quality novel-view synthesis by optimizing freely placed primitives in 3D and adaptively densifying them in under-reconstructed regions. However, this scene-adaptive capacity allocation is largely lost in existing feed-forward 3DGS methods, which commonly regress Gaussians at input pixels and lift them along camera rays. Such pixel-aligned formulations make the number and placement of primitives depend on image resolution and input viewpoints rather than scene complexity, resulting in dense and often redundant Gaussian sets. We present ATSplat, a feed-forward 3DGS framework that restores the adaptive allocation capability of 3DGS optimization through Adaptive 3D Tokens. ATSplat first lifts coarse patch-level depth and camera cues into sparse 3D anchor tokens, forming a compact scaffold of the scene. Each token is then regressed into local Gaussians with learnable 3D offsets, decoupling primitive placement from input image grids. An Adaptive Token Expansion module predicts a token-level uncertainty score, supervised by rendering error maps, and selectively expands high-uncertainty tokens through learnable expansion layers. This sparse-to-adaptive formulation enables ATSplat to concentrate primitives in challenging regions while maintaining a compact representation. Experiments on two representative datasets, RealEstate10K and DL3DV, show that ATSplat achieves state-of-the-art rendering quality while reducing the number of Gaussians by more than 5.7× compared with dense feed-forward 3DGS methods. From 12 input images at 512×960 resolution, ATSplat completes reconstruction in less than a second using a single commercial GPU, and renders high-quality novel views at 1136 FPS (512×960) with only 311K Gaussians.
In Cho, Jeonghwan Cho, Mijin Yoo +2
Yonsei University, South Korea · National University of Singapore, Singapore