LVS: Local View Synthesis from Relative Camera Pose by Reusing Previous Views
Authors: Qizhou Huo, Xuan Sun, Yongfei Guo, Zhipeng Wang, Yuanhao Gong
Organizations: Changchun Institute of Optics, Fine Mechanics and Physics, Chinese Academy of Sciences · University of Chinese Academy of Sciences, Beijing, China · Chinese Academy of Sciences, Beijing, China
Interactive scene exploration requires frequent view updates, although small camera motions preserve much of the visible content. Conventional 3D Gaussian Splatting nevertheless renders each target view, leaving this image overlap unexploited. Reusing rendered images offers an alternative. Geometric warping alone cannot recover newly exposed content and remains sensitive to depth errors. We propose a per-scene framework that replaces repeated scene rendering for nearby views with relative-pose-guided RGB-D image reuse. Geometric warping uses depth and relative pose to transport source content, while a lightweight multiscale network predicts RGB residuals to correct artifacts and infer missing appearance. Cached source features further reduce repeated computation. On GS-render, residual refinement improves PSNR by 0.72~dB over pure warping; evaluations on captured and rendered scenes demonstrate low query latency. This separation of scene rendering from local view updates supports responsive scene exploration, with potential applications in augmented and virtual reality.
Figures & tables
Figure 1: Local view synthesis. 3DGS renders each target view from the scene representation (top). Our method reuses a rendered RGB-D image through relative-pose-guided warping and RGB residual correction (bottom).
Figure 2: Framework overview. Depth and relative pose guide warping, while a multiscale decoder combines cached source features and geometric cues to predict an RGB residual. Adding this residual to the warped image yields the target prediction. Dashed red arrows indicate training supervision.
Figure 3: Qualitative comparison on Counter ( 3.5∘ ) and Kitchen ( 15∘ ). The first column shows target/source images; other columns show predictions and mean absolute RGB error maps. Scores are PSNR/SSIM. The dashed divider separates external baselines from our component comparisons.
Domain
Method
PSNR ↑
SSIM ↑
LPIPS ↓
Params (M)
Query (ms)
Cache
Peak
7-Scenes
CheapNVS
16.77
0.601
0.505
9.917
8.88
11.4
264.2
7-Scenes
AdaMPI32
20.47
0.701
0.378
18.946
5.60
403.8
7697.1
7-Scenes
3DGS
18.79
0.674
0.458
45.853
2.60
—
518.0
7-Scenes
Ours
20.74
0.721
0.395
0.110
0.95
14.4
95.8
GS-render
CheapNVS
20.33
0.622
0.329
9.917
8.75
11.4
264.2
GS-render
AdaMPI32
25.89
0.791
0.213
18.946
5.72
403.8
7697.1
Table 1: Test quality and query cost. Cache and peak memory are in MiB; 3DGS parameter counts refer to stored Gaussian scalars. Best values are bold.
Variant
PSNR ↑
SSIM ↑
LPIPS ↓
Params (K)
Warp-only
26.166
0.8340
0.1898
0.0
Network-only
19.659
0.5833
0.4299
109.7
Residual-1/4
26.526
0.8338
0.1910
104.5
Residual-1/2
26.706
0.8343
0.1902
108.7
Residual-Full
26.887
0.8407
0.1888
110.4
Table 2: Component ablation on GS-render, averaged over three scenes. Best values are bold.
Figure 4: Viewpoint sensitivity on Blender. PSNR (left) and LPIPS (right) are evaluated at matched target poses.
Per-scene 3D Gaussian Splatting (3DGS) enables high-fidelity rendering, but practical robotic and AR scene capture pipelines often depend on external geometric initialization (e.g., SfM point clouds or depth estimates), which can be slow and brittle in on-site deployment. We present ACEsplat, a fast per-scene optimization framework that reconstructs 3D Gaussian representations from RGB images and camera poses only, without requiring external 3D priors (e.g., precomputed SfM models or supervised depth maps). ACEsplat uses a two-stage pipeline: (1) a self-supervised scene coordinate regression (SCR) module builds an internal geometry prior within 4--5 minutes; (2) SCR features and coordinate priors are fused by a lightweight Gaussian initialization head, followed by per-scene 3DGS optimization. On static-view rendering, ACEsplat achieves 29.11 dB PSNR on Wayspots with real-time SLAM poses and 33.20 dB on Cambridge Landmarks with SfM-refined poses. On RealEstate10K sparse-view novel view synthesis, it achieves competitive image fidelity under a challenging 2-view setting. ACEsplat completes scene-specific SCR mapping and 3DGS reconstruction within 15--25 minutes on a single GPU, making it a practical RGB+pose-only solution for rapid scene setup in robotics and mixed-reality applications.
Mingkai Liu, Haohua Que, Dikai Fan +7
Peking University · University of Georgia · Infinity Robotics +3
Online novel view synthesis requires a model to reconstruct a scene causally from a stream of observations while keeping it renderable at every moment. We present ReCoSplat, an online feed-forward Gaussian Splatting model supporting both posed and unposed inputs, with or without camera intrinsics. While assembling local Gaussians with camera poses scales better than canonical-space prediction, stable training requires ground-truth poses, creating a distribution mismatch when predicted poses are used at inference. To address this, we introduce a Render-and-Compare (ReCo) module. ReCo renders the accumulated scene from the viewpoint of the incoming observation, comparing the render with the observation to produce a stable conditioning signal that helps bridge the mismatch. To support long sequences, we propose a hybrid KV-cache compression strategy combining early-layer truncation with chunk-level selective retention, reducing the KV cache size by over 90% for 100 or more frames. ReCoSplat achieves state-of-the-art performance among online methods while processing 256-view streams at an average input throughput of 45.1 FPS, with an end-of-stream throughput of 41.1 FPS on an RTX 6000 Ada GPU. Code and pretrained models are released at https://freemancheng.com/ReCoSplat .
Freeman Cheng, Botao Ye, Xueting Li +3
University of California, Merced · ETH Zürich · NVIDIA +2
We present ViewSplat, a view-adaptive 3D Gaussian splatting network for novel view synthesis from unposed images. While recent feed-forward 3D Gaussian splatting has significantly accelerated 3D scene reconstruction by bypassing per-scene optimization, a fundamental fidelity gap remains. We attribute this gap to the limited capacity of single-step feed-forward networks to regress static Gaussian primitives that satisfy all viewpoints. To address this limitation, we shift the paradigm from static primitive regression to view-adaptive splatting. Instead of a rigid Gaussian representation, our pipeline learns a view-adaptive latent representation. Specifically, ViewSplat initially predicts base Gaussian primitives alongside the weights of scene-conditioned View MLPs. During rendering, these MLPs take target-view coordinates as input and predict view-dependent residual updates for each Gaussian attribute (i.e., 3D position, scale, rotation, opacity, and color). This mechanism, which we term view-adaptive splatting, allows each primitive to rectify initial estimation errors, effectively capturing high-fidelity appearances. Extensive experiments demonstrate that ViewSplat achieves state-of-the-art fidelity while maintaining fast inference and real-time rendering; our large backbone variant runs at 15 FPS during inference and 90 FPS during rendering. Our project page is available at https://cvlab-uos.github.io/ViewSplat.
Moonyeon Jeong, Seunggi Min, Suhyeon Lee +1
University of Seoul · Korea Electronics Technology Institute