MoonGS: High-quality Representation of the Lunar Surface via Gaussian Splatting Using Robust Depth Features from Image Pairs
Authors: Yun Jiang, Bo Zheng, Yingying Zhang, Xueming Xiao, Tao Hu, Hutao Cui, Zhiguo Meng, Ke Gao, +2 more
Organizations: Jilin University, Changchun, China · Shanghai Aerospace Control Technology Institute, Shanghai, China · Beijing Institute of Control Engineering, Beijing, China · Changchun University of Science and Technology, Changchun, China · Harbin Institute of Technology, Heilongjiang, China
High-quality 3D reconstruction of lunar terrain from sparse rover images is indispensable for autonomous lunar exploration, but remains challenging because viewpoint overlap is insufficient, surface textures are weak, and data volume is limited. We propose MoonGS, the first feed-forward 3D Gaussian Splatting framework tailored to lunar scenes. Given only two input images, MoonGS predicts pixel-aligned Gaussian primitives in a single forward pass and renders photorealistic novel views without any per-scene optimization. MoonGS (i) adopts an adaptable backbone design that seamlessly integrates advanced vision foundation models to extract robust depth features; (ii) integrates semantic priors in two manners: merging semantic cues with visual features to refine Gaussian parameter estimation, and adopting a semantic ranking loss that regularizes background depth; and (iii) employs an entropy-guided heuristic resampling strategy to augment sparse observations by selecting the most informative distant viewpoints with negligible overhead. Experiments on the LuSNAR benchmark and our synthetic weak-texture MoonBlender dataset show that MoonGS surpasses state-of-the-art feed-forward NeRF/3DGS baselines by +4.9 dB PSNR, +0.29 SSIM, and 40% lower LPIPS while maintaining sub-second inference. Furthermore, we validate the broad applicability of our framework by demonstrating that it effectively leverages state-of-the-art backbones, including VGGT, to significantly boost performance. Qualitative evaluations on Chang'e mission imagery also show the best visual quality among compared methods, indicating robustness on real lunar data. The source code and dataset are publicly available at https://github.com/InRobots/MoonBlender.
Fig. 2: Overview of MoonGS. From a pair of input images and their associated semantic information, MoonGS performs forward inference to obtain the parameters of Gaussian primitives that represent the lunar surface scene. It then utilizes a camera with a known pose to unproject these Gaussian primitives onto the reliable depth positions. With this explicit representation, images can be rendered from camera viewpoints that have not been directly observed.
Fig. 3: Architecture diagram. (a) MoonGS utilizes DUSt3R to extract multi-level robust depth features. (b) Multi-level features are fused and refined through 2D U-Net. (c) The model constructs a cost volume using the extracted features and optionally integrates semantic embeddings as inputs for prediction to obtain depth maps and Gaussian parameters. (d) It then unprojects the Gaussian parameters into space, synthesizes novel views using Gaussian splatting, and employs MSE for supervised learning, with an optional ranking loss. (e) Finally, the model implements a resampling strategy to select additional positional images that can improve rendering quality as inputs. Dashed arrows denote optional components.
Fig. 4: Epipolar‐line visualization of three reconstruction challenges. Three types of epipolar lines illustrate difficulties encountered during reconstruction in lunar environments. Ref.1 and Ref.2 depict two viewpoints of the same lunar scene. Colored points in Ref.1 correspond to regions indicated by matching-colored epipolar lines in Ref.2. The blue lines denote regions with insufficient viewpoint overlap, resulting in invalid epipolar constraints. The red lines indicate weakly textured background regions, where epipolar constraints are prone to errors. The green lines mark areas with limited observations at a distance, making high-fidelity texture prediction challenging.
Fig. 5: Color distribution histograms for semantic categories. The color histograms demonstrate significant differences in color distribution between specific semantic categories, while the color distributions for the same semantic category show minimal variation. These patterns demonstrate that regions belonging to different semantic categories possess distinct color characteristics, whereas regions within the same semantic category exhibit high similarity. This property allows semantic information to serve as an effective feature for input.
Fig. 6: Gaussian coverage statistics. First, we count the number of valid pixels occupied by each Gaussian. Then, for each pixel, we sum the occupied pixel counts from all Gaussians that cover it. The entire process runs in parallel on the GPU.
Fig. 7: Semantic pixel distributions and real-data evidence. We show LuSNAR and MoonBlender distributions, the sky versus other pixel ratio for Chang’e imagery, and a real Chang’e example. We randomly sampled 96 CE4 TCAM images. 53 images had sky background, and the reported ratios are computed over this subset.
Methods
LuSNAR [ 55 ]
MoonBlender
Inference Time
Trainable Params
PSNR ↑
SSIM ↑
LPIPS ↓
PSNR ↑
SSIM ↑
LPIPS ↓
(s)
(M)
Du et al. [ 36 ]
13.26
0.347
0.727
14.01
0.374
0.722
2.870
-
MuRF [ 10 ]
14.32
0.283
0.443
14.88
0.287
0.441
-
-
pixelSplat [ 1 ]
18.89
0.525
0.360
25.26
0.689
0.166
0.483
118
MVSplat [ 2 ]
16.99
0.224
0.385
25.11
0.676
0.192
0.323
12.0
MoonGS(Ours)
23.80
0.816
0.134
26.08
0.721
0.139
0.200
18.1
TABLE I: Quantitative comparisons. We evaluated our method by rendering three novel view images from four reference viewpoints for each scene. The performance was determined by averaging across all scenes. The inference time includes both scene encoding and rendering time, tested on a single Q6000 GPU. The number of trainable parameters was computed with pytorch-lightning [ 61 ] .
Fig. 8: Qualitative results on real Chang’e imagery. Given reference image pairs (left column) to render an intermediate viewpoint at the interpolated camera pose. We show three groups (top to bottom) under distinct illumination conditions. MoonGS remains consistent with the input pairs and is more robust to lighting variation, producing fewer artifacts than pixelSplat and MVSplat. PixelSplat exhibits reconstruction failures and noticeable artifacts in challenging lighting, and MVSplat is less robust to illumination changes. Red boxes indicate regions of interest for comparison, and arrows in the reference pair mark the locations of these regions.
Fig. 9: Qualitative comparison of novel views on the LuSNAR (top) and MoonBlender (bottom) test sets. Compared to the baseline, our method is able to generate more accurate and perceptually attractive images, especially for lunar regolith close to the camera.
Fig. 10: Comparison of rendering results with and without heuristic resampling strategy. We present two groups of rendering results across four viewpoints: (a, c) with the heuristic resampling strategy, and (b, d) without the heuristic resampling strategy. The resampled region ( Position 2 ) is the selected resampling location. The results demonstrate that heuristic resampling effectively reduces artifacts and produces a more realistic textural appearance.
Fig. 11: Qualitative comparison of depth maps on the MoonBlender dataset. This figure compares depth maps produced by several models on the MoonBlender dataset. Our proposed method (MoonGS) generates both accurate and smoothly varying depth maps. In contrast, MVSplat converges to erroneous depths, and although pixelSplat produces relatively accurate depths, it exhibits noticeable artifacts—such as non-smooth transitions (indicated by the arrows). These findings highlight the importance of robust depth estimation techniques for lunar scene reconstruction.
Methods
PSNR ↑
SSIM ↑
LPIPS ↓
pixelSplat [ 1 ]
19.22
0.632
0.355
MVSplat [ 2 ]
17.80
0.542
0.377
MoonGS + DUSt3R [ 41 ]
23.91
0.823
0.130
MoonGS + VGGT [ 42 ]
24.04
0.830
0.128
TABLE II: Applicability analysis on the LuSNAR dataset. Evaluated at a resolution of 224×224 , the results show that our method effectively capitalizes on the VGGT backbone to outperform baselines, verifying its robust applicability and capacity to evolve with advanced foundation models.
Fig. 12: Qualitative comparison on the LUSNAR dataset. We evaluate the effectiveness of integrating the VGGT foundation model within our proposed framework. In terms of visual quality, our method outperforms the baseline model.
Fig. 13: Impact of ranking loss on depth estimation in weakly textured background areas. The results show the effect of ranking loss on improving depth estimation in regions with weak texture. In the absence of ranking loss, depth estimates in these areas are often imprecise due to the lack of texture. However, by introducing ranking loss, the model is able to refine the depth predictions in these challenging regions, leading to more reliable results. Moreover, integrating ranking loss with the pixelSplat model yields further improvements in depth accuracy. These results underscore the value of ranking loss in addressing depth estimation issues in weakly textured regions.
Fig. 14: Visualization of features under different ablation settings. The features extracted by the visual foundation model and their corresponding similarity maps in different ablation settings: w/o fusion, w/o U-Net, and base. Two points are selected (#1 in the sky region and #2 on the lunar regolith), and global cosine similarity is computed to assess feature compactness. The results demonstrate that, after full module processing, the features become more coherent and compact.
Methods
PSNR ↑
SSIM ↑
LPIPS ↓
LuSNAR [ 55 ]
MoonGS (Full)
23.80
0.816
0.134
w/o U-Net
23.29
0.786
0.137
w/o Semantic Fusion
23.02
0.781
0.141
MoonBlender
MoonGS (Full)
26.08
0.721
0.139
TABLE III: Ablation Studies. This table summarizes the results of removing semantic fusion and the 2D U-Net from our pipeline. Comparing the performance with and without these components highlights their essential role in producing high-quality novel view synthesis.
Fig. 15: t-SNE analysis of semantic fusion. The t-SNE analysis highlights the impact of semantic fusion on feature distribution. The left plot (w/o semantic) shows the feature distribution without semantic fusion, while the right plot (ours) presents the features after semantic fusion. Each color corresponds to a different semantic class. Incorporating semantic information leads to a more distinct separation between features.
Methods
PSNR ↑
SSIM ↑
LPIPS ↓
MoonGS
w/o Resample
22.46
0.767
0.151
Fixed Resample 1
22.76
0.795
0.186
Fixed Resample 4
23.97
0.848
0.162
Fixed Resample 7
23.88
0.850
0.157
Fixed Resample 11
23.38
0.812
0.175
TABLE IV: Ablation study of the heuristic resampling strategy. This table shows the impact of our heuristic resampling strategy on the performance of various models, evaluated after synthesizing novel views for each model. Superscript values denote the specific resampling position. Overall, our strategy enhances the final rendering quality.
3D Gaussian Splatting (3DGS) has emerged as an effective representation for novel view synthesis and 3D scene reconstruction, creating an increasing demand for reliable quality assessment. Unlike conventional image quality assessment (IQA), the quality of a 3DGS scene depends not only on the perceptual fidelity of rendered views, but also on scene-level factors such as spatial structure and cross-view consistency. Existing IQA methods are limited by their reliance on 2D perceptual cues, whereas general multimodal large language models (MLLMs) are not designed for stable quality regression and may produce unreliable judgments. To address these limitations, a multimodal quality assessment framework is developed for 3DGS scene understanding. First, a 3D-aware quality representation learning framework is introduced by augmenting a VGGT-based encoder with a dedicated quality head. Multi-view images are encoded into view-specific features and aggregated to capture cross-view consistency, while geometric cues are incorporated through joint modeling of depth and point-cloud-related structural information, enabling the learning of structure-aware quality representations beyond appearance-driven features. Second, a grounded multimodal reasoning mechanism is constructed by jointly feeding original images, depth maps, point cloud renderings, and camera parameters into a Qwen-based MLLM.
Jingxuan Su, Shenglin Wang, Tiesong Zhao +2
Guangdong Provincial Key Laboratory of Ultra High Definition Immersive Media Technology, School of Electronic and Computer Engineering, Peking University, Shenzhen 518055, China · Peng Cheng Laboratory, Shenzhen, 518055, China · College of Physics and Information Engineering, Fuzhou University, Fuzhou 350108, China
This paper introduces a fast method for high-quality 3D Gaussian Splatting (3DGS) reconstruction without traditional Structure-from-Motion (SfM). The proposed approach leverages 3D Foundation Models (3DFMs) for camera pose and point-cloud initialization, then jointly optimizes both camera poses and Gaussian primitives using a depth-guided loss function. This enables fast convergence even from rough initialization with as few as 50-60 input views. To further improve reconstruction quality in sparse-view scenarios, an MLP-based pose refinement module is introduced alongside depth-guided supervision from the foundation model. Extensive experiments on Mip-NeRF 360, Tanks and Temples, and RobustNeRF demonstrate that the proposed method achieves competitive reconstruction quality (23.61 dB PSNR, 0.19 LPIPS) while reducing training time to approximately three minutes per scene. The proposed method produces ready-to-use 3DGS models at a fraction of the time required by existing pipelines, making it suitable for near real-time applications in robotics, VR, and autonomous navigation.
Anurag Dalal, Daniel Hagen, Kjell G. Robbersmyr +1
Dep. of Eng. Sciences University of Agder Grimstad, Norway · Top Research Centre Mechatronics University of Agder Grimstad, Norway
Online 3D reconstruction from monocular image sequences is a challenging and ongoing research topic. 3D Gaussian Splatting (3DGS), leveraging its high-quality real-time rendering capability, empowers online 3D reconstruction to represent dense scenes with enhanced expressiveness, and thus holds great promise for a wide range of applications such as robotics and AR/VR. However, existing online 3DGS methods still suffer from some key challenges: fragile camera pose estimation due to the lack of global optimization, and low optimization efficiency in large-scale or long-sequence scenarios. To address these issues, we propose a robust and efficient online voxelized 3DGS reconstruction framework integrated with global Sim(3) optimization, which enables reliable camera tracking and efficient global loop closure for both camera poses and voxelized 3DGS. To accelerate the convergence of the voxelized 3DGS, we further introduce a color residual learning strategy, which not only boosts optimization speed but also enhances rendering quality. Extensive experiments on diverse indoor and outdoor datasets demonstrate that our method achieves state-of-the-art performance in both camera pose estimation accuracy and rendering quality, while retaining real-time efficiency. Additionally, we develop and deploy a real-world UAV-based active reconstruction system grounded on our proposed method, validating its robustness and generalizability for practical online 3D reconstruction tasks. Our code and data are available at https://github.com/TrickyGo/MoonSplat.
Guo Pu, Yixuan Han, Haofeng Li +3
Wangxuan Institute of Computer Technology, Peking University, China · Beijing Hydrogen Intelligent Tech. Co., Ltd., China