Real-world scenes are compositional: bricks, blades of grass, pebbles, and tree leaves recur across human-built and natural environments. Existing neural scene representations model these elements independently. Most 3D Gaussian Splatting and follow-up abstraction and compression methods treat each element as unique, fitting millions of independent Gaussians per scene. Prior methods like Splat and Replace fit template objects, but they require mostly manual selection of repeated elements. As a result, these representations store redundant parameters and provide weak manipulation handles for downstream tasks. We introduce SCION, a hierarchical compositional scene representation that replaces independent Gaussians with a compact vocabulary of reusable primitives and lightweight world-space instances that place transformed copies throughout the scene. We fit this representation to multi-view captures via a joint optimization over discrete and continuous scene parameters, combining two-level densification over splats and instances with an adversarial loss that preserves detail across shared primitives. The recovered structure yields a compact, controllable representation while maintaining high quality even at 1.2 MB. SCION achieves rate-distortion favorable to existing Gaussian compression methods, and it enables instance-level scene editing and animation without retraining. Our results show that neural scene representations need not memorize scenes as independent primitives; they can discover reusable parts. Project webpage: https://light.princeton.edu/SCION
Figures & tables
Figure 1: Overview of SCION . We initialize instances from a sparse point cloud, group them into reusable primitives, assemble the selected local splats for rendering, and update the representation through image losses, projected adversarial supervision, and active densification.
Figure 2: Neural primitive assembly. The (a) learned primitives are (b) instantiated multiple times throughout the scene, resulting in (c) assembled primitive structures that (d) construct the scene.
Method
PSNR ↑
SSIM ↑
LPIPS ↓
MB ↓
ContextGS
20.79
0.559
0.511
1.760
HAC++
26.51
0.756
0.299
1.699
OMG
26.25
0.770
0.271
1.353
Ours-Small
22.01
0.605
0.431
1.204
Table 1: Matched-storage comparison on MipNeRF360. Each baseline is rerun at the lowest available compression point point near our method.
Figure 3: Qualitative matched-storage comparison. At approximately 1.5 MB, SCION retains coherent repeated structure as each reused primitive is represented once and instantiated many times.
Ours-Small ( n=3 )
Ours-Large ( n=3 )
Ablation
PSNR ↑
SSIM ↑
LPIPS ↓
MB ↓
PSNR ↑
SSIM ↑
LPIPS ↓
MB ↓
Splat DC only
21.98±0.07
0.56±0.00
0.45±0.01
1.48
23.23±0.02
0.65±0.00
0.34±0.00
3.81
+ instance DC
21.87±0.15
0.56±0.01
0.46±0.02
1.45
23.20±0.05
0.65±0.00
0.34±0.00
3.76
+ LPIPS
21.90±0.08
0.56±0.00
0.45±0.01
1.45
23.09±0.07
0.65±0.00
0.33±0.00
3.76
Full ( + adv)
21.74±0.12
0.55±0.01
0.37±0.02
1.44
22.88±0.02
0.63±0.00
0.27±0.00
3.72
Table 2: Component ablation. Rows cumulatively add the active SCION components on the MipNeRF360 garden anchor scene. “DC” stands for density control.
Figure 4: Component ablation renders. Example render of the Ours-Large variant of our method on the Garden scene from MipNeRF360. “DC” stands for density control.
Figure 5: Animation. Flow maps for the same wind animation applied to (a) 3DGS and (b) SCION . Per-splat transformations in 3DGS produce incoherent local motion, while SCION moves repeated scene elements coherently through instance transforms.
Feed-forward 3D Gaussian Splatting methods reconstruct a scene from posed or pose-free images in a single forward pass, yet current approaches predict one Gaussian per input pixel, tying the representation budget to camera resolution rather than scene complexity. A flat wall and a richly textured object thus produce equally many Gaussians despite very different geometric needs. We propose ZipSplat, a token-based feed-forward model that decouples Gaussian placement from the pixel grid. A multi-view backbone extracts dense visual tokens, and k-means clustering compresses them into a compact set of scene tokens. Cross- and self-attention refine these tokens, and a lightweight MLP decodes each into a group of Gaussians with unconstrained 3D positions. Because clustering is applied at inference, a single trained model spans the quality-efficiency curve without retraining. ZipSplat operates without ground-truth poses or intrinsics, yet sets a new state of the art on DL3DV and RealEstate10K with ∼6× fewer Gaussians than pixel-aligned methods, surpassing the best pose-free baseline by 2.1dB and 1.2dB PSNR, respectively. It further generalizes zero-shot to Mip-NeRF360 and ScanNet++, outperforming all comparable baselines. Our project page is at https://veichta.com/zipsplat.
3D representations are fundamental to scene rendering, understanding, and interaction. Recent approaches, such as 3D Gaussian Splatting and Neural Radiance Fields, achieve impressive photorealistic novel-view synthesis, but lack the ability to easily decompose scene elements into a few primitives, requiring additional segmentation or grouping for object-level manipulation. We present MLP-Splatting, a method that enables scene decomposition via a few expressive light-field primitives while providing photorealistic novel-view synthesis. MLP-Splatting models each primitive as an independent compact MLP with localized spatial support that predicts radiance and opacity. In contrast to low-level Gaussian primitives or a single global radiance field, our neural primitives provide greater expressive capacity while remaining spatially localized. Rendering is performed through efficient sparse volumetric compositing over ray-primitive interactions. Our primitives are supervised using RGB supervision alone, which yields primitives that represent local scene regions often corresponding to objects or object parts, enabling interactive object-level editing without segmentation masks by selecting a handful of primitives. Our method, augmented with optional semantic feature distillation, enables open-vocabulary scene interaction and open-set instant segmentation. Compared to state-of-the-art methods, we achieve substantially lower memory usage (1/15×) and faster rendering (3×), as we show in our experiments compared to semantic 3DGS methods. Project Page: https://shinjeongkim.com/mlp-splatting
3D scene reconstruction, modeling, and rendering are highly relevant for numerous tasks, and 3D Gaussian splatting has become a standard choice in this context. Its feed-forward variants provide fast reconstruction from sparse input views but often produce per-pixel primitives, leading to highly redundant and thus inefficient representations. We present a structure-aware merging pipeline that takes per-pixel primitives from any feed-forward method and consolidates them into a compact, content-adaptive Gaussian set while largely retaining visual quality at just 201th of the Gaussians of a per-pixel method. We group spatially coherent Gaussians of similar appearance into variable-size clusters via adaptive superpixel segmentation guided by a saliency map, which allocates fine segments to textured regions and coarse segments to homogeneous areas. We compress each cluster into a compact latent representation through a learned encoder, then match and consolidate representations across views based on geometric overlap and feature similarity via a learned merger. A level-of-detail decoder then produces the final Gaussians at a controllable resolution, enabling a flexible quality-efficiency trade-off at inference. As a post-processing module, the pipeline is backbone-agnostic, leveraging the strengths of existing feed-forward methods. This leads to better and more robust quality than achieved by previous approaches that target a reduction in primitive count, while providing a highly compact representation, that can be rendered efficiently.
Tim-Felix Faasch, Jochen Kall, Cyrill Stachniss
Bosch Research Hildesheim, Germany · University of Bonn Bonn, Germany · Lamarr Institute for Machine Learning and Artificial Intelligence Bonn, Germany