Feed-forward novel view synthesis methods achieve strong generalization from posed multi-view inputs, but scaling them to large input view sets remains challenging. Transformer-based approaches that jointly process all input-view tokens incur rapidly increasing computation and memory as the number of views grows, while simple view subsampling discards potentially useful observations. We introduce Anchor-Integrated Multi-View Synthesis (AIMS), a scalable framework that decouples the number of available observations from the number of views processed by the global synthesis model. AIMS selects a fixed set of spatially distributed anchor views using farthest point sampling, groups nearby observations around each anchor, and uses a lightweight learnable integrator to fuse their information into enriched anchor representations. This allows additional observations to contribute to synthesis while keeping the downstream global view budget fixed. Evaluations on RealEstate10K and ScanNet demonstrate a favorable quality--efficiency trade-off against transformer-based and Gaussian-based baselines. AIMS achieves 29.41 dB and 17.73 dB PSNR on the two datasets, respectively, with rendering averaging 7.24 ms per view.
Figures & tables
Figure 1: Overview of AIMS. Coverage-aware grouping selects anchors and neighboring views. The View Group Encoder compresses each group’s ray-conditioned image tokens into compact scene tokens. A global decoder combines these with target ray tokens to synthesize novel views.
RealEstate10K
ScanNet
Encode
Render
Method
PSNR ↑
SSIM ↑
LPIPS ↓
PSNR ↑
SSIM ↑
LPIPS ↓
(ms/scene) ↓
(ms/view) ↓
DepthSplat
18.28
0.6780
0.3169
15.70
0.5528
0.5118
1043.04
9.84
MVSplat
–
–
–
–
–
–
–
–
LVSM (E–D)
18.72
0.5751
0.3896
12.69
0.4954
0.6916
758.93
9.62
LVSM (D)
24.32
0.8009
0.1770
17.34
0.5698
0.5496
–
1584.35
Efficient-LVSM
22.06
0.7301
0.2093
14.68
0.5453
0.6271
112.08
52.25
Table 1: Reconstruction quality and inference time for baselines evaluated on all observations selected by coverage-aware view grouping with G=12 and K=5 . Encoding is performed once per scene; rendering uses cached representations where available, while LVSM (D) reports full per-view inference. MVSplat fails due to out-of-memory errors. Best results are bold.
Figure 2: Qualitative comparison with transformer-based baselines, including LVSM encoder–decoder (E–D), LVSM decoder-only (D), and Efficient-LVSM, on RealEstate10K and ScanNet.
RealEstate10K
ScanNet
Encode
Render
Method
Anchors
PSNR ↑
SSIM ↑
LPIPS ↓
PSNR ↑
SSIM ↑
LPIPS ↓
(ms/scene) ↓
(ms/view) ↓
DepthSplat
2
22.97
0.8017
0.1980
11.36
0.4345
0.6177
41.66
7.21
4
24.03
0.8598
0.1526
13.55
0.5180
0.5754
65.15
7.30
8
22.21
0.8305
0.1834
15.21
0.5612
0.5223
114.10
8.50
12
21.19
0.8002
0.2108
15.76
0.5757
0.4971
172.59
9.07
MVSplat
2
21.50
0.7590
0.2196
8.52
0.1810
0.6437
34.08
4.11
Table 2: Reconstruction quality and inference time for baselines using only 2 , 4 , 8 , or 12 anchor views. AIMS retains G=12 anchors and K=5 neighbors per anchor. Encoding is performed once per scene; rendering uses cached representations where available, while LVSM (D) reports full per-view inference. Best results are bold.
Figure 3: Quality–cost trade-off across anchor counts for controlled 10-epoch training runs with G∈{4,8,12,16} and fixed K=5 . (a) Number of scene tokens, (b) reconstruction quality on RealEstate10K and ScanNet, and (c) scene encoding and per-view rendering time.
RealEstate10K
ScanNet
K
PSNR ↑
SSIM ↑
LPIPS ↓
PSNR ↑
SSIM ↑
LPIPS ↓
1
29.07
0.8916
0.0973
16.82
0.5649
0.5352
3
29.50
0.8988
0.0932
17.49
0.5749
0.5237
5
29.41
0.8971
0.0943
17.73
0.5774
0.5201
7
29.23
0.8941
0.0961
17.82
0.5779
0.5192
Table 3: Effect of inference-time neighborhood size K on reconstruction quality. All settings use the same AIMS model trained with G=12 and K=5 , without retraining. Best results are bold.
Figure 4: Analysis of VGE attention and its sensitivity to camera pose. (a) Head-averaged attention maps at VGE layer 6 for learnable queries 0, 64, and 128 across a local view group. (b) Cross-view attention agreement under camera-pose perturbations, measured by the mean Spearman rank correlation over matched physical locations across 50 ScanNet scenes.
Appendix figures & tables2 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 5: Additional qualitative comparisons with transformer-based baselines, including LVSM encoder–decoder (E–D), LVSM decoder-only (D), and Efficient-LVSM, on RealEstate10K and ScanNet.
Figure 6: More visualizations on head-averaged attention maps at VGE layer 6 for learnable queries 0, 64, and 128 across a local view group.
Recent Large View Synthesis Models (LVSMs) advocate an encoder-decoder architecture that separates reconstruction and rendering into distinct networks. We re-examine this design. Through controlled experiments, we show that a decoder-only architecture, which represents scenes implicitly as a KV-cache, outperforms encoder-decoder variants while using fewer parameters at identical rendering complexity. Further analysis shows that sharing weights between the color-input reconstruction network and the camera-only rendering network better aligns their features at the same viewpoint, facilitating image synthesis. Building on this finding, our model, dubbed DVSM, further incorporates foundation model priors and stage-wise patch sizing for an improved efficiency-quality tradeoff. Our results establish a new state of the art for novel-view synthesis across multiple benchmarks, in some cases even outperforming per-scene-optimized 3DGS under dense input views.
Large view synthesis models synthesize novel views through cross-view attention without explicit 3D representations, and recent studies have shown that they learn accurate spatial correspondence from RGB supervision alone. We observe that this correspondence generalizes beyond appearance. When non-photorealistic signals such as binary encoded panoptic labels are passed through the model, they are propagated to novel views with consistent spatial structure. These results indicate that the correspondence learned for RGB view synthesis can also propagate view-independent per-pixel labels. From this observation, we present the first work to extend large view synthesis models beyond appearance rendering to 3D scene understanding. We propose a panoptic segmentation pipeline that reuses a frozen view synthesis model to propagate panoptic labels from input views to novel views, without 3D reconstruction or any segmentation-specific training of the view synthesis model. Given panoptic labels on the input views, we encode them into binary channel representations and pass them through the same model to render target-view segmentation. On ScanNet, our method achieves segmentation quality on par with Gaussian based approaches requiring explicit 3D reconstruction, while outperforming them in novel view synthesis by more than 7 dB. The label propagation also transfers across datasets, surpassing these approaches on Replica without any fine-tuning.
Recently, novel view synthesis has witnessed remarkable progress, with mainstream methods such as Neural Radiance Fields (NeRF) and 3D Gaussian Splatting (3DGS) delivering impressive results. However, these approaches often struggle to balance rendering speed and model size, and their optimization-based training can be highly time-consuming. Furthermore, they typically rely on dense observations, often failing to produce satisfactory results under sparse-view conditions. Although feed-forward reconstruction significantly reduces the optimization time of 3DGS, its pixel-aligned formulation generates millions of Gaussians from a single image, severely limiting its practical deployment on mobile devices. To address these limitations, we revisit the Multiplane Image(MPI) representation, which represents scenes using a compact set of planar layers for efficient novel view synthesis. Leveraging recent advances in visual foundation models, we utilize predicted point maps for reliable geometric initialization, followed by differentiable optimization. To address the issues of holes and artifacts in sparsely initialized MPI, we introduce one-step diffusion, which participates in both the differentiable optimization of MPI and the postprocessing of rendering results. Compared with a representative GS-based method, our approach is 30.7% faster and uses only 14.8% of its model size, while achieving competitive synthesis quality on front-view scenarios