PlenoCI: Plenoptic CharacterIstics for View Dependence Aware Change Classification
Authors: Jason Lai, Chamuditha Jayanga Galappaththige, Niko Suenderhauf, Dimity Miller, Donald G. Dansereau
Organizations: Australian Centre for Robotics, School of Aerospace, Mechanical and Mechatronic Engineering, The University of Sydney · QUT Centre for Robotics · ARIAM Hub
Radiance field representations such as 3D Gaussian Splatting (3DGS) natively encode complex visual phenomena such as occlusions and view dependence, but they are inherently underconstrained. Independently optimized reconstructions converge to different primitive configurations, even in unchanged regions. We introduce Plenoptic CharacterIstics (PlenoCI), a novel feature built from the plenoptic field these representations approximate. PlenoCI directly captures rich visual behaviors while ignoring Lambertian textures. By deriving closed-form analytic plenoptic derivatives from a 3DGS representation, we efficiently detect these 5D structures. Our approach is robust to underconstrained representations by construction, reporting two orders of magnitude fewer false positives between independent reconstructions of unchanged scenes than concurrent work. We demonstrate PlenoCI's utility on change classification. First, we detect changes with an instance-aware 3DGS pipeline, achieving state-of-the-art results on CL-Splats with a 25.7% mIoU gain over the strongest competitor, while remaining competitive on the more challenging PASLCD benchmark. Leveraging PlenoCI, we classify changes as geometric or appearance-based with a balanced accuracy of 0.735, comparable to the best performing baseline. We believe plenoptic derivatives and PlenoCI open new directions for view dependence aware understanding in visually complex environments. Code and data are available at https://js0n-lai.github.io/plenoci.
Figures & tables
Figure 1 : Given two 3DGS reconstructions from different times, our method first detects scene changes. To classify changes, we draw on view-dependent plenoptic structures encoded in our novel PlenoCI feature. For example, we correctly detect the sheet of paper and book as changes, and classify them using PlenoCI as occlusion edge evidence (left). Our approach is robust to the inherently underconstrained nature of 3DGS representations; two independent optimizations given the same training images yield different primitive configurations. Primitive-space comparisons [ 22 ] report false positives in static scenes while our method does not (right).
Figure 2 : Example scenario in flatland to motivate our novel PlenoCI feature. Given a 2D world ( Fig. 2(a) ), we build the 3D plenoptic field L and compute partial derivatives L∗ ( Figs. 2(b) to 2(c) ). These are useful signals for detecting visually complex phenomena such as occlusion edges ( Fig. 2(d) ). Brightness represents magnitude corresponding to each color channel’s plenoptic field. While we show ground truth geometry and L here, this also generalizes to radiance field representations.
Figure 3 : Our PlenoCI feature extraction pipeline. Given a 3DGS representation G , we sample analytic plenoptic derivatives ∇L over a local light field to compute the structure tensor. The third eigenvalue λ2 provides an informative signal for complex visual behaviors where existing features are unreliable, while ignoring Lambertian textures. We mask pixelwise λ2 with an edge image (dilated for clarity) to yield PlenoCI Rk from view k . Repeating this over selected views yields R.
Figure 4 : Our overall scene change detection and classification approach. Given instance-aware 3DGS scene representations G0 and G1 , we aggregate per-ray change scores δ in the corresponding primitives. This yields a sparse set of anchors ΔGt′ to propagate on changed instances. For change detection, we sample edge rays Re to find changed Gaussians ΔGtall . For change classification, we leverage PlenoCI R to find geometrically changed Gaussians ΔGtgeo . These explicit representations can be used to render novel change masks with disambiguation between geometric and appearance changes.
Method
CL-Splats [ 1 ]
PASLCD [ 20 ]
mIoU ↑
F1 ↑
mIoU ↑
F1 ↑
CYWS [ 60 ]
0.495
0.640
0.273
0.398
GeSCF [ 35 ]
0.692
0.793
0.477
0.611
SceneDiff [ 75 ]
0.334
0.433
0.473
–
3DGS-CD [ 47 ]
0.700
0.792
0.209
0.339
MV3DCD [ 20 ]
0.634
0.752
0.478
0.628
Table 1 : Binary SCD results averaged over CL-Splats [ 1 ] and PASLCD [ 20 ] . PASLCD baselines are sourced from [ 23 , 75 ] . Our method achieves state-of-the-art performance on CL-Splats, and is competitive with the strongest PASLCD baselines. The first , second , and third best performances are highlighted.
Figure 5 : Qualitative comparison of binary change masks for CL-Splats [ 1 ] (rows 1–2) and PASLCD [ 20 ] (rows 3–4). Pairwise approaches such as GeSCF [ 35 ] are susceptible to spurious detections. Our instance-aware 3DGS approach outperforms all baselines on CL-Splats, and is competitive to state-of-the-art on the challenging PASLCD benchmark.
Method
mIoU ↑
F1 ↑
No RGB
0.530
0.671
No Semantics
0.469
0.614
No Occupancy
0.514
0.658
Ours
0.538
0.688
Table 2 : Change scoring ablation averaged over PASLCD [ 20 ] . Incorporating all elements provides the best performance.
Method
mIoU ↑
F1 ↑
Precision ↑
Recall ↑
Ours
0.538
0.688
0.712
0.691
Oracle
0.596
0.731
0.667
0.839
Table 3 : SCD results averaged over PASLCD [ 20 ] varying instance segmentations. Oracle derives instances from ground truth masks while we use learned instances [ 43 ] . Recall is the main metric, improving by 20% with oracle instances as the learned instances undersegment small changed objects.
Figure 6 : Qualitative change classification results for PASLCD [ 20 ] . Top to bottom: Porch, Lounge, Meeting Room, Zen. Our PlenoCI feature enables building explicit geometry and appearance-based change representations for rendering multi-class change masks.
Metric
GS-Diff [ 22 ]
Ours
Oracle
Balanced Accuracy ↑
0.792
0.735
0.851
Geo Precision ↑
0.970
0.985
0.970
Geo Recall ↑
0.896
0.745
0.871
App Precision ↑
0.478
0.431
0.463
App Recall ↑
0.687
0.725
0.830
Table 4 : CC performance averaged over PASLCD [ 20 ] . The best performance excluding Oracle is bolded . Oracle uses instances derived from ground truth masks ( Sec. 5.2 ), driving improvements in recall of both classes. Ground truth class imbalance favoring Geo amplifies App classification errors.
Method
Binary (%) ↓
Geo (%) ↓
App (%) ↓
GS-Diff [ 22 ]
0.389
0.064
0.325
Ours
0.004
0.001
0.003
Table 5 : False positive rate from comparing two independently optimized 3DGS models of an unchanged scene, averaged over PASLCD [ 20 ] . Scenes with real changes contain 3.51% change pixels on average, ranging from 0.17–20.12%.
Figure 7 : Pixelwise eigenvalue signals derived from the plenoptic structure tensor. Given an RGB view (column 1), we sample analytic plenoptic derivatives over a local light field to build pixelwise structure tensors M . The eigenvalues λi of M provide informative signals of scene structures, such as edges (column 2), corners (column 3) and occlusion edges as well as specular highlights (column 4).
Figure 8 : Propagation of change information from anchors ΔGt′ (column 1, highlighted in green) to produce dense change representations ΔGt (column 3, highlighted in green) depends on the quality of the instance-aware 3DGS (column 2). While our anchors correctly identify changed objects, upstream undersegmentation (rows 1–2) can cause our instance gating to attenuate these anchors, creating false negative predictions (column 4). With oracle segmentations (rows 3–4), our predicted changes are considerably more accurate.
Metric
All Rays
Edge Rays (Ours)
mIoU ↑
0.517
0.538
F1 ↑
0.661
0.688
Ext TFLOPs ↓
35.0
34.9
Ext VRAM (GB) ↓
10.3
5.13
Sco GFLOPs ↓
1.11
0.043
Sco VRAM (GB) ↓
3.64
1.92
Table 6 : Comparison of ray selection approaches averaged over all scenes in PASLCD [ 20 ] using 20 sample views each. Floating point operations reported for Extraction (Ext) do not include rasterization, which is shared across both methods. Retaining only edge rays (Ours) marginally improves SCD performance, but substantially reduces compute costs for change scoring (Sco).
Condition
Ours (%) ↓
GS-Diff (%) ↓
View Overlap
100% (full vs. full)
0.000
0.487
50% (half vs. half)
0.003
0.397
0% (half vs. half)
0.020
0.446
Coverage
Full vs. full
0.000
0.487
Table 7 : False positive rate from comparing two independently optimized 3DGS reconstructions of the Cantina scene from PASLCD [ 20 ] under no changes. Under varying training view overlaps and levels of coverage, our method consistently reports fewer false positives than GS-Diff [ 22 ] .
Figure 9 : Comparison of independently optimized 3DGS models of the Cantina scene from PASLCD without changes. At 100% view overlap, both models receive all training views. GS-Diff [ 22 ] reports a false positive rate of 0.487% as it directly compares primitives (inset), while we correctly report no changes. At 25% view coverage, Model B receives 5 out of 20 training views. Poor reconstruction quality causes our approach to detect false positives, albeit at a lower rate than GS-Diff.
Scene change detection methods built on Gaussian splatting universally follow a render-then-compare paradigm: the pre-change scene is rendered into 2D and compared against post-change images via pixel or feature residuals. This change detection problem with Gaussian Splatting has been treated as a question about pixels; we treat it as a question about primitives. We provide direct evidence that native primitive attributes alone -- position, anisotropic covariance, and color -- carry sufficient signal for scene change detection. What makes primitive-space comparison hard is the under-constrained nature of Gaussian splatting representation: independent optimizations yield primitive solutions whose count, positions, shapes, and colors differ even where nothing has changed. We address this challenge with anisotropic models of geometric and photometric drift, complemented by a per-primitive observability term that reflects the extent to which each Gaussian is constrained by the camera geometry. Operating directly on primitives gives our method, GD-DIFF, two properties that distinguish it from render-then-compare methods. First, change maps are multi-view consistent by construction, where prior work had to learn this through an additional optimization objective. Second, geometric and appearance changes are scored separately, identifying not just where but what kind of change occurred, distinguishing structural changes (e.g., an added object) from surface-level ones (e.g., a color change) without supervision or external model dependencies. On real-world benchmarks, GS-DIFF surpasses the prior state-of-the-art approach by ∼17% in mean Intersection over Union.
Chamuditha Jayanga Galappaththige, Jason Lai, Timothy Patten +3
1QUT Centre for Robotics · 2ARIAM · 3ACFR, University of Sydney +1
Factories, museums and surveyors photograph the same space months apart and need to know which objects changed. When each visit is reconstructed with 3D Gaussian Splatting (3DGS), a direct comparison of the two reconstructions does not answer this. Training is stochastic, so two reconstructions of an unchanged space never coincide, and the second visit is often a quick re-scan with far fewer photographs. We propose GS-Pool, which takes two independently reconstructed Gaussian fields of the same space and returns the changed objects in each, together with their masks. SAM2 masks of each visit's photographs are lifted onto the Gaussians that render them and merged into an object pool, so every decision is taken once per object in 3D. We introduce a photographic carrier, the 3DGS training loss of each input reconstruction against the other visit's photographs, backpropagated to the Gaussians that rendered each pixel. We combine it with GS-Diff's geometry and colour terms and our distilled DINOv3 features. This evidence is compared with that of the objects present in both visits, which sets a change threshold for each scene. On PASLCD, GS-Pool reaches mIoU/F1 scores of 0.751/0.846 against 0.644/0.758 for GS-Diff, the strongest prior method, a gain of 17%/12%. Its mIoU is also 36%, 40% and 57% above that of O-SCD, PlenoCI and MV-3DCD, and it reaches 0.855 mIoU on CL-Splats, 33% above MV-3DCD. Each changed object is returned as a set of Gaussians with the evidence behind its decision, which an inspector can review in 3D.
3D Gaussian Splatting (3DGS) provides an explicit and efficient scene representation, but its primitives lack inherent object-level identity, hindering downstream tasks such as open-vocabulary scene understanding. Existing methods typically address this by either distilling high-dimensional feature embeddings into Gaussians or by lifting 2D mask labels into 3D via heuristic refinement. However, feature-based approaches incur heavy storage and decoding overhead, while lifting-based pipelines remain vulnerable to label contamination: Gaussians necessary for appearance reconstruction often receive incorrect object labels during 2D-to-3D projection. We propose OP2GS, an object-aware Gaussian representation that augments each primitive with an explicit instance identity and a dedicated instance opacity σ∗ for object-mask rendering. The original opacity σ remains responsible for visual reconstruction, while σ∗ models whether a Gaussian should contribute to a particular object mask. This dual-opacity formulation decouples visual existence from instance occupancy: mislabeled Gaussians can remain available for image rendering while becoming transparent in the object-mask branch. To learn this representation, we introduce a random object loss that optimizes the 1D instance occupancy field using the standard transmittance-based visibility of 3DGS. Semantic descriptors are then attached at the object level through multi-view aggregation, eliminating per-Gaussian feature storage. Compared with feature-training approaches, OP2GS achieves competitive open-vocabulary performance while significantly reducing computational overhead. Compared with training-free pipelines, it leverages physically consistent occupancy learning to resolve visibility ambiguities.
Guiyu Liu, Niklas Vaara, Janne Mustaniemi +2
Center for Machine Vision and Signal Analysis, University of Oulu, Finland · Aalto University, Finland