Organizations: School of Information Technology and Engineering (SITE), Kazakh-British Technical University, Almaty, Kazakhstan · School of Information Technology and Engineering (SITE), ADA University, Baku, Azerbaijan
Accurate assessment of color differences is essential for applications ranging from digital design to quality control. While existing color difference metrics, such as CIEDE2000, aim to approximate human perception, they may still exhibit inconsistencies with perceptual judgments. In this study, we investigate a data-driven approach to color-difference estimation based directly on human evaluations. We collect similarity judgments for 2,000 systematically generated color pairs, each rated by seven observers using a four-point ordinal scale. These judgments are then used to train regression models using different color representations, including RGB channel differences, HSI differences, and COLIBRI fuzzy linguistic categories. Experiments with five regression algorithms show that the choice of color model has a greater influence on prediction performance than the choice of regression algorithm. Using COLIBRI features alone, linear regression achieves an R2 of 0.595, outperforming RGB and HSI representations, which achieve R2 values of 0.479 and 0.493, respectively. The best performance is obtained by LightGBM using the combined representation, reaching an R2 of 0.703. The results indicate that human perceptual color differences are better captured when numerical color coordinates are complemented by graded perceptual categories, highlighting the potential of data-driven models for perceptually aligned color-difference estimation.
Figures & tables
Study
Separation conditions
Color pairs
Color difference range
Assessment method
Key finding (separation effect)
Vision science foundations
Boynton et al., 1977 [ 7 ]
Juxtaposed; with gap
—
Threshold
Threshold detection
Foundational gap effect study: gap impaired luminance discrimination but improved chromatic discrimination; explained by contour enhancement and spatial averaging
Sharpe & Wyszecki, 1976 [ 8 ]
Juxtaposed; with gap
—
Threshold & supra
Ratio comparison; liminal det.; color matching
Separation impairs lightness more than chromaticness; proposed a proximity factor for color-difference formulas
Eskew, 1989 [ 9 ]
Juxtaposed; narrow gap (isoluminant fill)
—
Threshold
Threshold detection
Luminance or chromatic gap prevents spatial integration, enhancing chromatic sensitivity; gap had little effect for flashed stimuli
Danilova & Mollon, 2006 [ 10 ]
Juxtaposed; separated up to 10°
—
Threshold
2AFC threshold
Discrimination optimal at small gap; thresholds rose moderately with increasing separation; gap effect more pronounced in parafovea
Table I : Summary of Prior Studies on the Effect of Sample Separation on color-Difference Evaluation. Studies are grouped by stimulus type. The number of observers ranged from 4 to 46 across studies; see original sources for details. NS = no separation; — = not reported in original source.
Figure 1 : Color pairs visualization
Figure 2 : Interfaces used in the color similarity judgement experiments
Figure 3 : Proposed framework for perceptual color difference modeling
Dataset characteristic
Value
Participants
7
Unique color pairs
2 000
Total evaluations
28 600
Human similarity rating distribution
Not similar
9 501 (33.22%)
Somewhat similar
7 192 (25.15%)
Table II : Summary of the perceptual color similarity dataset and grouped cross-validation design.
Metric
Feature Set
Linear Regression
Decision Tree
Random Forest
LightGBM
XGBoost
RMSE
RGB
0.2653±0.003
0.2794±0.006
0.2519±0.004
0.2438±0.004
0.2530±0.005
HSI
0.2616±0.004
0.2546±0.007
0.2312±0.003
0.2268±0.001
0.2344±0.002
COLIBRI
0.2337±0.004
0.2654±0.009
0.2341±0.006
0.2274±0.005
0.2364±0.003
RGB + HSI
0.2552±0.004
0.2409±0.005
0.2155±0.003
0.2116±0.003
0.2161±0.003
RGB + COLIBRI
0.2248±0.005
0.2422±0.003
0.2118±0.001
0.2033±0.002
0.2091±0.002
HSI + COLIBRI
0.2315±0.004
0.2413±0.002
0.2143±0.002
0.2104±0.002
0.2156±0.001
Table III : Performance comparison of regression models across different feature sets using five-fold grouped cross-validation. Values are shown as mean ± standard deviation across folds.
Model Dependency
Feature
Coefficient (Mean ± SD)
RGB
HSI
COLIBRI
RGB + HSI
RGB + COLIBRI
HSI + COLIBRI
RGB + HSI + COLIBRI
Intercept
0.6930±0.005
0.6792±0.005
0.7851±0.005
0.6954±0.005
0.7899±0.005
0.7799±0.005
0.7937±0.005
Red Channel
−0.0027±0.0
-
-
−0.0014±0.0
−0.0015±0.0
-
−0.0018±0.0
RGB
Green Channel
−0.0029±0.0
-
-
−0.0016±0.0
−0.0013±0.0
-
−0.0016±0.0
Blue Channel
−0.0024±0.0
-
-
−0.0013±0.0
−0.0011±0.0
-
−0.0014±0.000
Hue
-
−0.7848±0.007
-
−0.4954±0.010
-
−0.2282±0.008
0.0950±0.016
Table IV : Linear regression coefficients for perceptual color similarity across different feature representations using five-fold grouped cross-validation. Values are shown as mean ± standard deviation across folds.
Figure 4 : Experimental results for Experiment 1 and 2 side to side
Category
Experiment 1
Experiment 2
Very Similar
ΔE≤4.221
ΔE≤2.069
Similar
4.221<ΔE≤11.406
2.069<ΔE≤6.743
Somewhat Similar
11.406<ΔE≤24.266
6.743<ΔE≤18.889
Not Similar
ΔE>24.266
ΔE>18.889
Table V : Perceptual similarity decision boundaries thresholds
Figure 5 : Comparison of the proposed model with conventional color-difference measures.
Pair
Model Difference
Model Similarity
Model Interpretation
ΔE00
CIEDE2000 Interpretation
1–2
0.539
0.461
Somewhat similar
18.066
Very much
3–4
1.000
0.000
Not similar
68.840
Strongly
5–6
0.982
0.018
Not similar
74.104
Strongly
7–8
0.403
0.597
Similar
3.047
Appreciable
9–10
0.444
0.556
Similar
25.478
Strongly
11–12
0.498
0.502
Similar
9.713
Much
Table VI : Linguistic comparison between the proposed model and CIEDE2000.
Evaluating the perceptual alignment between Contrastive Vision-Language Models (CVLMs) and humans is typically constrained by traditional benchmarks that overlook fine-grained semantic and cultural nuances. In this work, we propose a novel evaluation framework that leverages the gamified, discrete color space of the board game Hues and Cues. By mapping the board's 480 color cells to the CIE xy chromaticity diagram, we calculate empirical perceptual distances across a carefully curated 100-word vocabulary spanning seven semantic categories. To properly contextualize model performance, we establish an empirical lower bound of expected error-the Human Consistency baseline-calculated via Leave-One-Out (LOO) cross-validation on a dense dataset of color associations collected from 325 human observers through a custom digital interface. We evaluate 162 models across multiple architectural families and pre-training datasets to assess their semantic color grounding. Our results demonstrate that while CVLMs successfully replicate human cognitive biases, such as idealized memory colors for concrete physical referents (e.g., food and plants), they systematically diverge from the human baseline in abstract, subjective, and pop-culture domains. We identify two distinct failure modes in severely misaligned concepts: semantic misclassification and a systematic uncertainty collapse into a default blue coordinate. Furthermore, we reveal that highly curated pre-training datasets are significantly more effective than massive, uncurated corpora in mitigating these severe misalignments. Ultimately, this work highlights that despite their broad categorization capabilities, current CVLMs still fail to capture the nuanced, localized consensus of human color memory, emphasizing the value of gamified tasks in exposing underlying model biases. The data and code are publicly available to test other metrics.
Nuria Alabau-Bosque, Jorge Vila-Tomás, Paula Daudén-Oliver +3
Image Processing Lab, Universitat de València, Carrer del Catedrátic José Beltrán Martinez, Paterna, 46980, Spain
Do vision models see colors the way humans do? Existing evaluations of color representations usually compare them with geometric spaces such as CIELAB or with discrete color labels. These references capture perceptual distance or category membership, but not the graded way in which people organize colors. We evaluate color grounding against a fuzzy perceptual model with 86 graded categories fitted to human survey data. The framework can be applied to any image encoder and measures three complementary properties: category boundaries, category compactness, and graded alignment beyond what color geometry alone can explain. Across eleven Vision Transformer encoders, the category-level results are broadly similar, whereas graded alignment differs substantially. Masked Autoencoders achieve the strongest beyond-geometry alignment, with confidence intervals that do not overlap those of the other encoders. A layer-wise analysis further shows that masked reconstruction preserves this structure toward the output. On natural images, MAE represents surface color globally, while language-supervised models encode color more strongly in relation to the foreground object. These results show that human-like color grounding has several distinct aspects that should not be reduced to a single score.
Ayan Igali, Pakizar Shamoi
School of Information Technology and Engineering Kazakh–British Technical University Almaty, Kazakhstan
Human color categories are not uniformly distributed in perceptual space, yet most computational color models still assume fixed and evenly structured representations. In this paper, we present a focused analytical extension of the COLIBRI fuzzy color model by investigating perceptual asymmetry between hue categories. Using previously collected large-scale human color categorization data, we introduce quantitative measures of category extent and boundary uncertainty, namely Wideness and Boundary Width, derived from fuzzy membership functions at the α = 0.5 level. The analysis reveals a strong imbalance between the two categories: yellow occupies a compact and sharply constrained region of the hue space, whereas green spans a substantially broader interval and exhibits a more extended transition structure. The results show that perceptual color categories are not only fuzzy, but also highly non-uniform in their geometric organization. This asymmetry suggests that some categories behave as narrow, highly specific perceptual labels, while others function as broad, tolerant regions of human color naming. These findings provide a new perspective on linguistic color categorization and extend the interpretability of the COLIBRI framework for perceptually grounded color modeling.