Open-vocabulary 3D scene understanding enables object localization and segmentation from free-form text queries without a fixed category vocabulary. Many recent methods build on 3D Gaussian Splatting and consolidate multi-view observations, such as masked crops from individual views, into language features or compact object descriptors before the query is known. However, observations of the same object vary across viewpoints and are not equally informative: some reveal cues relevant to a particular query, whereas others provide incomplete or misleading evidence. Pre-query consolidation can therefore suppress cues on which a later query depends. We introduce EviSplat, which preserves individual observation features as evidence for later text queries. EviSplat retains individual observation features within class-agnostic 3D instances that represent objects, object parts, or background regions. It also learns, for each Gaussian, a distribution describing which visual appearances its observations support. Given a text query, EviSplat scores each instance using its most relevant observations. It then computes a score for each Gaussian by combining instance-level relevance with locally supported evidence, weighted by how often and how unambiguously that Gaussian was observed. Different queries can thus draw on different visual cues from the same preserved evidence. Experiments across diverse datasets and evaluation protocols demonstrate state-of-the-art performance, supporting the benefit of preserving multi-view evidence until query time and aggregating it according to the query.
Figures & tables
Figure 1: Pre-query consolidation vs. EviSplat: preserve evidence. Top: features aggregated across observations are scored against the query. Bottom: EviSplat retains individual observation features for query-dependent selection and score aggregation. Crops illustrate the source observations; feature vectors and relevance bars are schematic. For the query “jake” in the same view, the masks show LaGa (top, IoU 0.46 ) and the full EviSplat model (bottom, IoU 0.91 ).
Figure 2: Overview of EviSplat. Top: multi-view evidence preservation, illustrated for one instance i (three of its retained observations shown) and one Gaussian g seen in three views. Bottom: query-guided evidence aggregation for the query “jake”. Histograms show a subset of the K prototype bins, colored consistently with the codebook; only the latents and decoders (purple) are trained. The two stages share the observation banks Bi , codebook C , and learned distributions pg .
Table 1: LERF-OVS segmentation. † denotes our reproduced baseline for LaGa ( Cen et al., 2025b ) ; other prior results are cited as reported value. Metrics are four-scene averages, and our EviSplat reports three-run means ± population SD. mAcc: mask IoU >0.25 success rate ( Best / Second-Best ).
School of Electrical Engineering and Robotics, Queensland University of Technology (QUT), Brisbane, QLD 4000, Australia · CSIRO Robotics, CSIRO, Brisbane, QLD 4069, Australia