cs.CVSep 23, 2026

Overlapping Visual Grouping Without Semantic Priors

Authors: Teemu Saukkio, Hashem Haghbayan, Juha Plosila

Organizations: University of Turku, Faculty of Technology, Department of Computing, Turku, Finland

Abstract

Most computer-vision systems organize visual input toward a predefined interpretation, such as semantic categories, prompted regions, learned object-like representations, or a single spatial partition. This work considers an earlier stage of visual organization: the formation of candidate perceptual units directly from sensor measurements before their identity, meaning, or task relevance is known. We introduce Domain Parent Grouping (DPG), a sensor-grounded grouping method in which complementary measurement relationships are represented in separate processing domains. Spatially connected groups formed within these domains are related through cross-domain overlap, yielding a non-exclusive grouping representation rather than a single mutually exclusive segmentation. This representation retains broader and more localized groups, as well as alternative grouping boundaries over the same image locations, simultaneously available. DPG also includes a native mechanism for reprocessing selected group content, in which input-relative measurement ranges allow the observational resolution to change while preserving previously formed groups. DPG is implemented using three domains representing locally contextualized luminance, direct chromatic relationships, and contextual chromatic relationships. Experiments on the BSDS500 dataset demonstrate the benefit of combining the three domains. The results further show that DPG forms measurement-supported groups corresponding to low-level image structure, and that these groups exhibit measurable correspondence with human-annotated regions and boundaries. This demonstrates that structured visual organization can emerge directly from relationships among sensor measurements.

Figures & tables

Appendix figures & tables2 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. More Accurate, Less Human: Gestalt Grouping in Vision Models

    Aug 10, 2026Sudhanva Manjunath Athreya, Sai Phani Kumar MalladiVision Foundation Models

  2. CRISP: Compositional Relations as Invariant Structural Priors for Domain Generalization

    May 7, 2026Dat Nguyen, Duc-Duy NguyenMultimodal Domain GeneralizationFine-Grained Perception

  3. Learning visual representations for compositional analysis of artworks and photographs

    Aug 6, 2026Fatemeh Behrad, Tinne Tuytelaars, Johan WagemansVisual RepresentationsArt