Most computer-vision systems organize visual input toward a predefined interpretation, such as semantic categories, prompted regions, learned object-like representations, or a single spatial partition. This work considers an earlier stage of visual organization: the formation of candidate perceptual units directly from sensor measurements before their identity, meaning, or task relevance is known. We introduce Domain Parent Grouping (DPG), a sensor-grounded grouping method in which complementary measurement relationships are represented in separate processing domains. Spatially connected groups formed within these domains are related through cross-domain overlap, yielding a non-exclusive grouping representation rather than a single mutually exclusive segmentation. This representation retains broader and more localized groups, as well as alternative grouping boundaries over the same image locations, simultaneously available. DPG also includes a native mechanism for reprocessing selected group content, in which input-relative measurement ranges allow the observational resolution to change while preserving previously formed groups. DPG is implemented using three domains representing locally contextualized luminance, direct chromatic relationships, and contextual chromatic relationships. Experiments on the BSDS500 dataset demonstrate the benefit of combining the three domains. The results further show that DPG forms measurement-supported groups corresponding to low-level image structure, and that these groups exhibit measurable correspondence with human-annotated regions and boundaries. This demonstrates that structured visual organization can emerge directly from relationships among sensor measurements.
Figures & tables
Figure 1: Example from BSDS500. The original test image 100007 is shown on the left and visualization of the human-annotated segmentation is shown on the right. The annotation represents the principal mutually exclusive regions but does not enumerate all potentially overlapping or illumination-dependent visual structures.
Figure 2: Qualitative comparison for the brighter capture, GOPR1789. The top row shows the mutually exclusive partitions produced by the three comparison methods. The bottom row shows the simplified visualization of the complete DPG group set and the original input image.
Figure 3: Parent-child organization associated with the red bucket. The parent retains most of the bucket within one group despite internal measurement variation, while its child groups represent more localized regions.
Figure 4: Example of grouping at multiple levels of spatial extent. The displayed groups were selected after processing based on their spatial correspondence with the tree structures. The original processing produces both a larger tree-associated parent group and a more spatially restricted group behind the trampoline. Reprocessing the restricted group exposes additional internal organization at a finer observation level.
Figure 5: A grouping hypothesis associated with the illuminated foreground structure. The group extends across spatially detailed vegetation and does not correspond to a single semantic object or a conventional mutually exclusive scene segment.
Figure 6: Qualitative comparison for the lower-luminance capture, GOPR1792. The top row shows the mutually exclusive partitions produced by the three comparison methods. The bottom row shows the simplified visualization of the complete DPG group set and the original input image.
Figure 7: Reprocessing of a broad parent group under lower-luminance conditions. The initial parent contains substantial internal variation. Reprocessing makes more spatially restricted parent groups associated with the tree and red bucket separately available.
Figure 8: Black-container-associated parent obtained by reprocessing the broad parent in Figure 7 . Despite weak separation from parts of the surrounding low-luminance vegetation, most of the visible container extent forms a distinct group in the restricted observation context.
Figure 9: Comparison of the individual DPG domains, their pairwise combinations, and the complete three-domain DPG method over the 200 BSDS500 test images. Each bar shows the mean of the 200 image-level values for the corresponding metric, and the error bars show the sample standard deviation across those images. L denotes the luminance domain, VC the vector-colour domain, and CC the colour-consistency domain. DPG denotes the complete combination of all three domains.
Figure 10: Comparison of the complete DPG method with Felzenszwalb segmentation, watershed with boundary-RAG merging, and Quickshift over the 200 BSDS500 test images. Each bar shows the mean of the 200 image-level values for the corresponding metric, and the error bars show the sample standard deviation across those images.
Figure 11: Illustrative training-set example using image 12003 and region 2 from annotation 1. The top row shows the human-annotated region within the original image, the region separately, and the maximum-IoU DPG proposal. The bottom row shows the maximum-IoU proposals selected independently for the comparison methods. The IoU and B-F1 values shown in the panels are calculated only between the selected human region and the displayed proposal; they are not the image-level metrics used in the 200-image test-set evaluation. The DPG proposal has the highest IoU with the selected reference region in this illustrative example.
Figure 12: Illustrative training-set example using image 108073 and region 3 from annotation 1. The top row shows the human-annotated region within the original image, the region separately, and the maximum-IoU DPG proposal. The bottom row shows the maximum-IoU proposals selected independently for the comparison methods. The IoU and B-F1 values shown in the panels are selected-region/proposal-pair measures and are not the image-level metrics used in the 200-image test-set evaluation. The selected DPG proposal follows the directly illuminated part of the tiger within the annotated region.
Figure 13: Illustrative training-set example using image 118035 and region 2 from annotation 1. The top row shows the human-annotated region within the original image, the region separately, and the maximum-IoU DPG proposal. The bottom row shows the maximum-IoU proposals selected independently for the comparison methods. The IoU and B-F1 values shown in the panels are selected-region/proposal-pair measures and are not the image-level metrics used in the 200-image test-set evaluation. The DPG proposal contains the complete white building as one group.
Appendix figures & tables2 assets
Supplementary material from the paper’s appendix.
Appendix
Metric
Comparison
Complete DPG mean ± SD
Comparison mean ± SD
Δ
95% CI
pH
rrb
Covering
L
0.5251 ± 0.1095
0.3644 ± 0.1021
0.1607
[0.1484, 0.1737]
8.61688×10−34
0.994
Covering
VC
0.5251 ± 0.1095
0.4146 ± 0.1191
0.1104
[0.1003, 0.1212]
8.61688×10−34
1.000
Covering
CC
0.5251 ± 0.1095
0.3888 ± 0.1229
0.1363
[0.1240, 0.1489]
8.61688×10−34
1.000
Covering
L+VC
0.5251 ± 0.1095
0.4678 ± 0.1120
0.0573
[0.0497, 0.0654]
8.61688×10−34
1.000
Covering
L+CC
0.5251 ± 0.1095
0.4469 ± 0.1122
0.0782
[0.0692, 0.0874]
8.61688×10−34
1.000
Covering
VC+CC
0.5251 ± 0.1095
0.4621 ± 0.1187
0.0629
[0.0552, 0.0711]
8.61688×10−34
1.000
Appendix
Table 1: Paired comparisons between complete DPG and the six component configurations. For each metric, Δ is the mean paired image-level difference defined as complete DPG minus the comparison configuration. CI denotes the 95% paired-bootstrap confidence interval, pH the within-metric Holm-adjusted Wilcoxon p -value, and rrb the matched-pairs rank-biserial correlation.
Metric
Comparison
Complete DPG mean ± SD
Comparison mean ± SD
Δ
95% CI
pH
rrb
Covering
Felzenszwalb
0.5251 ± 0.1095
0.5260 ± 0.1255
-0.0010
[-0.0167, 0.0141]
0.77151
0.024
Covering
Watershed–RAG
0.5251 ± 0.1095
0.4238 ± 0.1493
0.1013
[0.0818, 0.1211]
1.06555×10−17
0.705
Covering
Quickshift
0.5251 ± 0.1095
0.2171 ± 0.0689
0.3079
[0.2896, 0.3265]
4.30844×10−34
1.000
Recall@0.50
Felzenszwalb
0.2729 ± 0.1190
0.2674 ± 0.1293
0.0054
[-0.0107, 0.0213]
0.336721
0.081
Recall@0.50
Watershed–RAG
0.2729 ± 0.1190
0.2247 ± 0.1111
0.0481
[0.0311, 0.0653]
3.74230×10−7
0.432
Recall@0.50
Quickshift
0.2729 ± 0.1190
0.1327 ± 0.0900
0.1402
[0.1228, 0.1581]
1.62001×10−29
0.932
Appendix
Table 2: Paired comparisons between complete DPG and the three conventional grouping methods. For each metric, Δ is the mean paired image-level difference defined as complete DPG minus the comparison method. CI denotes the 95% paired-bootstrap confidence interval, pH the within-metric Holm-adjusted Wilcoxon p -value, and rrb the matched-pairs rank-biserial correlation.
Human vision organizes what it sees into wholes: same-colored points group into series, similar marks cohere into categories, and shapes complete into recognizable objects. These are the Gestalt operations that visualization design builds on. Whether vision models organize visual content this way has not been systematically tested. We introduce a behavioral battery that scores models against human data from prior perception studies on four grouping tasks: mark-color odd-one-out, color-series counting, silhouette recognition, and object odd-one-out. We apply it to 45 models across five training families: supervised, self-supervised, and contrastive vision-language encoders, open-weight VLMs, and closed foundation models. The battery reveals that agreement with human responses captures aspects of perceptual organization that conventional performance metrics fail to distinguish, with several closed models exhibiting substantially lower alignment than their benchmark accuracy would suggest. Scoring against published perception data therefore gives visualization research a reusable yardstick, requiring no new user study, for auditing whether the models now entering visualization pipelines organize what they see the way their human audience does.
Sudhanva Manjunath Athreya, Sai Phani Kumar Malladi
Domain generalization requires identifying stable representations that support reliable classification across domains. Domains may differ in low-level attributes, such as color, texture, or visual style, while preserving the same structural relationships among their underlying components. Existing methods primarily address these differences by improving the training process or aligning features across domains. However, since they leave this shared compositional structure implicit, they may overlook a more reliable source of invariance and consequently generalize less effectively to unseen domains. We propose Compositional Relational Invariance from Spatial Primitives (CRISP), an image classification framework that factors visual recognition into visual primitives and their relational composition. We represent these compositions using soft unary, binary, and ternary predicates over primitive locations and appearance, yielding differentiable measures of spatial and visual alignment that can be learned end-to-end. To learn primitives and relational structure jointly, we design an end-to-end architecture with three components: (1) a visual backbone that extracts generalized features, (2) a concept bottleneck layer that maps these features to primitive heatmaps with differentiable spatial coordinates, and (3) a structural scoring layer that evaluates candidate spatial relations among the detected primitives. Finally, we compute class probability from the joint evidence of its class-specific relational compositions and localized primitive appearance. We evaluate \method{} on five real-world image-classification datasets from the widely used DomainBed suite, covering shifts in depiction style, dataset provenance, and camera-trap location and achieving the new state-of-the-art on both benchmarks.
Dat Nguyen, Duc-Duy Nguyen
Harvard University Basis Research Institute · Hanoi University of Science and Technology
Composition, the deliberate arrangement of visual elements, is central to how meaning, emotion, and aesthetic quality are conveyed in artwork, yet it remains among the least formalized dimensions of visual understanding. Prior work highlights a persistent gap in learning meaningful compositional representations, attributing it to semantic bias and suggesting that human-inspired approaches may be key. We compare two parallel paradigms for composition analysis: a human-inspired method grounded in perceptual grouping, and fine-tuned foundation models enabled by recent large-scale compositional datasets. The human-inspired approach uses object-centric models for region-level decomposition and a graph attention network to capture spatial relationships between elements. Both paradigms are evaluated on composition score/category prediction, compositional image retrieval, and visual saliency detection. With frozen encoders, the human-inspired method achieves competitive performance while remaining interpretable. When sufficient data enables fine-tuning, large self-supervised models outperform significantly, but at the cost of interpretability and cross-domain generalization.