Factories, museums and surveyors photograph the same space months apart and need to know which objects changed. When each visit is reconstructed with 3D Gaussian Splatting (3DGS), a direct comparison of the two reconstructions does not answer this. Training is stochastic, so two reconstructions of an unchanged space never coincide, and the second visit is often a quick re-scan with far fewer photographs. We propose GS-Pool, which takes two independently reconstructed Gaussian fields of the same space and returns the changed objects in each, together with their masks. SAM2 masks of each visit's photographs are lifted onto the Gaussians that render them and merged into an object pool, so every decision is taken once per object in 3D. We introduce a photographic carrier, the 3DGS training loss of each input reconstruction against the other visit's photographs, backpropagated to the Gaussians that rendered each pixel. We combine it with GS-Diff's geometry and colour terms and our distilled DINOv3 features. This evidence is compared with that of the objects present in both visits, which sets a change threshold for each scene. On PASLCD, GS-Pool reaches mIoU/F1 scores of 0.751/0.846 against 0.644/0.758 for GS-Diff, the strongest prior method, a gain of 17%/12%. Its mIoU is also 36%, 40% and 57% above that of O-SCD, PlenoCI and MV-3DCD, and it reaches 0.855 mIoU on CL-Splats, 33% above MV-3DCD. Each changed object is returned as a set of Gaussians with the evidence behind its decision, which an inspector can review in 3D.
Figures & tables
Figure 2 : GS-Pool on one PASLCD cell: a distilled DINOv3 field per visit (1), lifted SAM2 segments merged into objects (2), four carriers per primitive and the scene’s null (3), one decision per object (4), and the change mask and changed primitives (5).
Figure 3 : A glass emptied in place on Lounge (a). Geometry and colour rank it near the scene’s middle (b, c), semantics higher (d) and the photographs near the top (e). (f) GS-Pool ’s mask and the annotation.
mean
mIoU per reconstruction
F1 per reconstruction
Method
Frame
Annotations
mIoU
F1
seed 22
23
24
seed 22
23
24
PASLCD
GS-Pool (FastPGSR)
joint
PASLCD
0.7509
0.8461
0.7404
0.7520
0.7602
0.8379
0.8466
0.8537
GS-Pool (vanilla 3DGS)
joint
PASLCD
0.7198
0.8209
0.7223
0.7111
0.7259
0.8236
0.8144
0.8248
GS-Pool (vanilla 3DGS)
anchored
PASLCD
0.7281
0.8266
0.7368
0.7139
0.7335
0.8341
0.8159
0.8300
GS-Pool (FastPGSR)
anchored
PASLCD
0.7260
0.8277
0.7135
0.7347
0.7300
0.8171
0.8351
0.8309
Table 1: Change detection on PASLCD and CL-Splats. GS-Pool is given as the seed mean and per seed. Baselines are as reported ( MV-3DCD on PASLCD as reported in [ 22 ] ), except the rows on our CL-Splats annotations, which we ran ( † four of five scenes).
Figure 4 : Porch at the test view with the most annotated change: the annotation, GS-Pool on FastPGSR and on vanilla 3DGS, and O-SCD , each tagged with its view IoU.
Figure 5 : The output in 3D on Porch from a viewpoint no camera took. Each visit’s changed primitives are shown on its own reconstruction and the rest in grey: T1 ’s two stools and mug (a), and T2 ’s stool and the objects on its table (b).
FastPGSR
vanilla 3DGS
row
decision
mIoU
F1
mIoU
F1
1
photographs alone
0.4681
0.5896
0.4353
0.5627
2
+ semantics
0.5899
0.7087
0.6060
0.7291
3
all four carriers, averaged
0.6090
0.7252
0.6107
0.7357
4
core, F(o)>θ
0.7034
0.8080
0.6617
0.7762
5
+ photo rank, quorum
0.7113
0.8147
0.6668
0.7796
Table 2: The decision rule built up one component at a time on the same pools and carriers. Mean over the 60 joint-frame members per trainer. Row 6 is the full rule.
Appendix figures & tables10 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 7 : SAM2’s stability threshold on the joint frame, mean over 60 members per trainer. Dashed: the chosen 0.90 . Dotted: the library’s default 0.95 .
row
prims/visit
time
FastPGSR, joint
0.19M
2.1 min
FastPGSR, anchored
0.17M
1.7 min
vanilla 3DGS, joint
1.14M
5.7 min
vanilla 3DGS, anchored
1.13M
4.2 min
Appendix
Table 3 : Median wall-clock per member on one RTX 4090 (24 GB), from the reconstructions, distilled fields, cached features and segmentation maps to the masks and change PLYs. Producing those inputs is excluded. Stage 3’s neighbour searches run on the GPU.
mIoU
F1
GS-Pool , re-scored by this pass
0.7509
0.8461
oracle: greedy decisions, same pool
0.8448
0.9090
oracle: every false pixel removed
0.8172
0.8882
oracle: every missed pixel filled
0.9229
0.9568
Appendix
Table 4: The decision oracle and the two pixel bounds on the same 60 members (FastPGSR, joint frame). The first row re-scores GS-Pool with the same pass and reproduces its score exactly.
FastPGSR
vanilla 3DGS
comparison
mIoU
F1
mIoU
F1
full rule
0.7509
0.8461
0.7198
0.8209
core alone (row 4)
0.7034
0.8080
0.6617
0.7762
null
full rule, own-visit null
0.7464
0.8407
0.7269
0.8205
full rule, both visits’ null
0.7357
0.8301
0.7195
0.8170
core alone, own-visit null
0.6530
0.7647
0.5986
0.7109
Appendix
Table 5: Twelve controlled comparisons around the full rule, each changing one thing, over the 120 joint members.
Method
mean
median
scenes below MV-3DCD
MV-3DCD [ 16 ]
1.72
1.34
—
O-SCD [ 22 ] †
1.47
1.09
3 of 9
GS-Pool (FastPGSR)
1.03
0.31
6 of 10
GS-Pool (vanilla 3DGS)
1.15
0.17
7 of 10
Appendix
Table 6 : Pixels flagged between two captures of the same unchanged scene under different lighting (%), at instance 2’s reference cameras ( † nine of ten scenes).
Method
mean
median
runs at zero
GS-Diff [ 4 ] ∗
0.389
—
—
PlenoCI [ 3 ] ∗
0.004
—
—
GS-Pool (FastPGSR)
0.153
0.017
14 of 60
GS-Pool (vanilla 3DGS)
0.511
0.017
15 of 60
Appendix
Table 7: Pixels flagged as changed between two reconstructions of an unchanged scene (%). ∗ As reported by PlenoCI [ 3 ] on its own reconstructions, not measured on our pairs.
Method
annotations
mIoU
F1
GS-Diff [ 4 ]
PlenoCI
0.724
0.829
O-SCD [ 22 ]
PlenoCI
0.756
0.845
MV-3DCD [ 16 ]
PlenoCI
0.634
0.752
PlenoCI [ 3 ]
PlenoCI
0.951
0.975
MV-3DCD [ 16 ]
GS-Pool
0.6413
0.7667
O-SCD [ 22 ] †
GS-Pool
0.6224
0.7428
Appendix
Table 8 : Change detection on the CL-Splats real scenes, PlenoCI’s figures on its own annotations [ 3 ] above and ours below ( † four scenes, 25 views).
Figure 10 : All five CL-Splats scenes at the view with the most annotated change, both trainers at seed 22, labelled with the view’s IoU and the scene’s three-seed mIoU (FastPGSR / vanilla 3DGS).
Figure 12 : Cells 1–10 (Cantina to Meeting_room) at the test view with the most annotated change, joint frame, seed 22, labelled with the view’s IoU and the cell’s three-seed mIoU (FastPGSR / vanilla 3DGS).
Figure 13 : Cells 11–20 (Playground to Zen), as in Figure 12 .
Scene change detection methods built on Gaussian splatting universally follow a render-then-compare paradigm: the pre-change scene is rendered into 2D and compared against post-change images via pixel or feature residuals. This change detection problem with Gaussian Splatting has been treated as a question about pixels; we treat it as a question about primitives. We provide direct evidence that native primitive attributes alone -- position, anisotropic covariance, and color -- carry sufficient signal for scene change detection. What makes primitive-space comparison hard is the under-constrained nature of Gaussian splatting representation: independent optimizations yield primitive solutions whose count, positions, shapes, and colors differ even where nothing has changed. We address this challenge with anisotropic models of geometric and photometric drift, complemented by a per-primitive observability term that reflects the extent to which each Gaussian is constrained by the camera geometry. Operating directly on primitives gives our method, GD-DIFF, two properties that distinguish it from render-then-compare methods. First, change maps are multi-view consistent by construction, where prior work had to learn this through an additional optimization objective. Second, geometric and appearance changes are scored separately, identifying not just where but what kind of change occurred, distinguishing structural changes (e.g., an added object) from surface-level ones (e.g., a color change) without supervision or external model dependencies. On real-world benchmarks, GS-DIFF surpasses the prior state-of-the-art approach by ∼17% in mean Intersection over Union.
Chamuditha Jayanga Galappaththige, Jason Lai, Timothy Patten +3
1QUT Centre for Robotics · 2ARIAM · 3ACFR, University of Sydney +1
Dynamic 3D Gaussian Splatting (3DGS) methods reconstruct time-varying scenes from synchronized multi-camera video using photometric supervision. When a moving object becomes fully occluded from all training cameras, this supervision vanishes: the Gaussians representing it receive no gradient signal and degrade. Existing approaches to incomplete observations in neural reconstruction rely on learned generative priors that prioritize visual plausibility over physical correctness. We propose PersistGS, a method that restores object permanence during occlusion by coupling differentiable rigid body simulation with 3D Gaussian Splatting. Our approach decomposes the scene into per-object Gaussians and collision meshes, estimates friction and velocity from the observed pre-occlusion trajectory via differentiable simulation, and uses the resulting SE(3) trajectory to position object Gaussians throughout the occlusion period. Because the predicted trajectory satisfies the governing equations of rigid body dynamics, it faithfully captures contact events (bounces, friction-based deceleration, direction changes) that kinematic extrapolation cannot model. We introduce a centroid silhouette loss that isolates positional gradients from appearance noise, yielding 40% lower trajectory error than photometric supervision. We evaluate using cameras withheld from training that observe the object during its occlusion. Experiments on synthetic scenes show that PersistGS outperforms constant velocity extrapolation by +2.46dB PSNR and comes within 0.19dB of a ground-truth trajectory upper bound.
3D Gaussian Splatting (3DGS) provides an explicit and efficient scene representation, but its primitives lack inherent object-level identity, hindering downstream tasks such as open-vocabulary scene understanding. Existing methods typically address this by either distilling high-dimensional feature embeddings into Gaussians or by lifting 2D mask labels into 3D via heuristic refinement. However, feature-based approaches incur heavy storage and decoding overhead, while lifting-based pipelines remain vulnerable to label contamination: Gaussians necessary for appearance reconstruction often receive incorrect object labels during 2D-to-3D projection. We propose OP2GS, an object-aware Gaussian representation that augments each primitive with an explicit instance identity and a dedicated instance opacity σ∗ for object-mask rendering. The original opacity σ remains responsible for visual reconstruction, while σ∗ models whether a Gaussian should contribute to a particular object mask. This dual-opacity formulation decouples visual existence from instance occupancy: mislabeled Gaussians can remain available for image rendering while becoming transparent in the object-mask branch. To learn this representation, we introduce a random object loss that optimizes the 1D instance occupancy field using the standard transmittance-based visibility of 3DGS. Semantic descriptors are then attached at the object level through multi-view aggregation, eliminating per-Gaussian feature storage. Compared with feature-training approaches, OP2GS achieves competitive open-vocabulary performance while significantly reducing computational overhead. Compared with training-free pipelines, it leverages physically consistent occupancy learning to resolve visibility ambiguities.
Guiyu Liu, Niklas Vaara, Janne Mustaniemi +2
Center for Machine Vision and Signal Analysis, University of Oulu, Finland · Aalto University, Finland