Modern vision models increasingly rely on rich semantic representations that extend beyond class labels to include descriptive concepts, attributes, and contextual cues. However, semantic concepts are not uniformly represented across classes: a concept may be frequent globally yet remain underrepresented within a class, resulting in low class-concept coverage. We formalize this class-concept phenomenon as Semantic Coverage Imbalance (SCI) and study its relationship with learned semantic representations. We find that lower class-concept coverage is consistently associated with weaker semantic representations. To address this, we introduce Coverage-aware Semantic Representation Learning, which shares more semantic structure for low coverage or uncertain relations while preserving more pair-specific structure for high coverage relations. We instantiate this principle in SemCovNet with Coverage-Calibrated Semantic Sharing (CCSS), a coverage- and uncertainty-aware partial-pooling mechanism. Using a range of vision datasets spanning facial attributes and fine-grained recognition, as well as real-world medical datasets, we show that SCI is widespread, degrades representation quality, and can be mitigated through coverage-aware semantic representation learning.
Figures & tables
Figure 1: Motivation and principle of coverage-aware semantic representation learning. (a) Semantic Coverage Imbalance ( SCI ) occurs when class–concept pair coverage is highly uneven, even when class and concept marginals appear adequate. (b) We show that low coverage weakens semantic representations. (c) Coverage-aware semantic representation learning adapts more pair-specific structure for high coverage pairs and shares more structure for low coverage or uncertain pairs.
Figure 2: Intuitive view of Semantic Coverage Imbalance ( SCI ). (a) Each class–concept pair (y,k) defines a relation in the semantic space S . (b) In the observation space X , the training data grounds each semantic relation as G:(y,k)↦Dy,k , producing high- and low coverage pairs. (c) We hypothesize that low coverage leads to weaker learned semantic representations. We refer to the phenomenon of uneven class–concept pair coverage as SCI .
Figure 3: SemCovNet for coverage-aware semantic representation learning. A Concept Query Module maps dense features to image-driven concept tokens. For each class–concept pair, CCSS combines the empirical class–concept representation, a shared semantic structure, and calibration signals. low coverage or uncertain pairs shrink toward the shared structure, while high coverage pairs retain more pair-specific structure. The resulting target supervises the concept tokens, and a residual pathway preserves complementary visual information for task prediction.
Figure 4: SCI across visual domains. Ranked training-set class–concept coverage Cy,k for CelebA, CUB, and Derm7pt. Coverage is highly uneven: 26.9%, 67.2%, and 33.9% of class–concept pairs have Cy,k<0.05 in CelebA, CUB, and Derm7pt, respectively.
Figure 5: Lower coverage is associated with weaker semantic representations. (a) Frozen linear-probe class-conditioned concept Macro-F1 from the learned concept tokens across Tail, Middle, and Head class–concept pairs for DINOv3–ERM and concept supervised learning. Error bars show 95% CIs over class–concept pairs. (b–d) Concept supervised representations projected using a common PCA basis fitted separately for each dataset. Dispersion varies across coverage groups and datasets; + denotes class–concept centroids.
Variant
Concept supervised
CCSS
Adaptation
CUB
CelebA
Task ↑
Tail ↑
ΔH−T↓
Task ↑
Tail ↑
ΔH−T↓
ERM
×
×
Frozen
0.886 ± 0.006
N/A
N/A
0.953 ± 0.001
N/A
N/A
Concept
✓
×
Frozen
0.881 ± 0.007
0.003 ± 0.002
0.171 ± 0.032
0.954 ± 0.004
0.441 ± 0.102
0.426 ± 0.091
Concept + CCSS (Ours)
✓
✓
Frozen
0.888 ± 0.002
0.005 ± 0.001
0.145 ± 0.011
0.956 ± 0.002
0.463 ± 0.006
0.406 ± 0.008
ΔTail (Concept+CCSS − Concept), Frozen [95% CI]
0.002 [0.001, 0.002]
0.022 [0.015, 0.060]
Concept
✓
×
LoRA
0.886 ± 0.003
0.004 ± 0.002
0.188 ± 0.026
0.949 ± 0.006
0.411 ± 0.072
0.448 ± 0.049
Table 1: Controlled CCSS model ladder on CUB and CelebA. The ladder isolates the effect of CCSS under frozen and LoRA-adapted DINOv3 encoders (last three layers). Concept denotes standard semantic supervision without CCSS. Tail reports class-conditioned concept Macro-F1 and ΔH−T the Head–Tail representation gap. Highlighted rows show the paired Tail gain from adding CCSS under the same adaptation setting, with 95% CIs.
Method
CelebA
CUB
Derm7pt
MILK10k
JTT
0.887 ± 0.004
0.890 ± 0.004
0.852 ± 0.008
0.879 ± 0.019
DFR
0.933 ± 0.001
0.837 ± 0.002
0.866 ± 0.011
0.859 ± 0.017
MORE
0.925 ± 0.005
0.892 ± 0.001
0.854 ± 0.009
0.879 ± 0.013
TS-MOF
0.907 ± 0.013
0.885 ± 0.001
0.858 ± 0.003
0.878 ± 0.014
IVQ-CBM
0.932 ± 0.002
0.881 ± 0.008
0.870 ± 0.005
0.861 ± 0.015
SemCovNet (Ours)
0.955 ± 0.001
0.891 ± 0.002
0.851 ± 0.018
0.887 ± 0.009
Table 2: Task-level comparison with existing methods. We report the standard task metric for each dataset: CelebA BAcc, CUB Top-1 accuracy, and AUROC for Derm7pt and MILK10k. Higher is better.
Figure 6: Controlled class–concept coverage intervention on CUB. (a) Frozen-probe representation quality as training evidence for target class–concept pairs is reduced from 100% to 0% ; “one” retains one positive example. (b) Paired zero-coverage effect Δ0 positive values indicate higher representation quality for SemCovNet . Error bars show paired 95% CIs over target pairs. At zero coverage, SemCovNet retains higher representation quality than Coverage reweight and Semantic GroupDRO.
Figure 7: Coverage-aware learning under clinical semantic supervision. (a) On Derm7pt expert-defined clinical concepts, SemCovNet improves Tail representation quality and reduces the Head–Tail gap over Concept supervision. (b) On MILK10k continuous soft semantic supervision, Tail representation quality remains stable as semantic evidence becomes less informative. Error bars show 95% CIs.
Appendix figures & tables13 assets
Supplementary material from the paper’s appendix.
Appendix
Setting
Configuration
Visual encoder
DINOv3 ViT-B/16
Input resolution
224 × 224
Encoder adaptation
Frozen / LoRA on final 3 transformer blocks
LoRA rank r
8
LoRA scaling α
16
LoRA dropout
0.05
Appendix
Table 3: Implementation details. Methods use the same backbone, data splits, optimization settings, and random seeds.
Dataset
#train
Class
Concept
Pairs
Q25
Median
Q75
Near-zero (%)
Zero (%)
CelebA
162770
2
39
78
0.048
0.143
0.338
26.9
0.0
CUB
5394
200
312
62,400
0.000
0.000
0.111
67.2
53.9
Derm7pt
413
2
28
56
0.031
0.117
0.455
33.9
12.5
Waterbirds
4795
2
2
4
0.050
0.500
0.950
25.0
0.0
MILK10k
3668
11
7
77
0.253
0.368
0.433
0.0
0.0
Appendix
Table 4: Dataset-level characterization of Semantic Coverage Imbalance. Coverage Cy,k is computed per class–concept pair on the training split only. Near-zero and Zero denote the fraction of pairs with Cy,k<0.05 and Cy,k=0 , respectively.
Dataset
Significant heterogeneity (%)
ρbal
CelebA
97.4
1.000
CUB
96.5
0.996
Derm7pt
64.3
0.993
Waterbirds
100.0
0.800
MILK10k
85.7
—
Appendix
Table 5: Statistical controls for SCI . Significant heterogeneity is the share of concepts whose across-class coverage variation exceeds a prevalence-preserving permutation null after BH–FDR correction ( q<0.05 ). ρbal is the Spearman correlation between full-data and class-balanced coverage estimates; “—” denotes not applicable.
Method
100%
75%
50%
25%
One
0%
Δ100%→0%
Concept supervised
0.675 [0.630, 0.717]
0.666 [0.619, 0.708]
0.661 [0.617, 0.703]
0.665 [0.619, 0.709]
0.661 [0.617, 0.704]
0.665 [0.619, 0.708]
0.010
Coverage reweighting
0.650 [0.608, 0.691]
0.657 [0.613, 0.698]
0.659 [0.618, 0.699]
0.661 [0.617, 0.702]
0.661 [0.620, 0.701]
0.661 [0.623, 0.697]
-0.012
Semantic GroupDRO
0.655 [0.614, 0.696]
0.658 [0.617, 0.698]
0.649 [0.605, 0.689]
0.659 [0.618, 0.698]
0.665 [0.621, 0.707]
0.660 [0.617, 0.698]
-0.005
LoRA Concept
0.664 [0.623, 0.704]
0.658 [0.618, 0.697]
0.665 [0.628, 0.702]
0.669 [0.628, 0.709]
0.661 [0.620, 0.703]
0.670 [0.629, 0.710]
-0.006
SemCovNet (Ours)
0.679 [0.631, 0.724]
0.672 [0.623, 0.717]
0.677 [0.629, 0.722]
0.676 [0.627, 0.721]
0.678 [0.631, 0.721]
0.682 [0.634, 0.726]
-0.003
Appendix
Table 6: Full coverage-response intervention on CUB. We report frozen-probe representation quality qrepr for each method as target-pair training evidence decreases from 100% to zero. “One” retains a single positive example and therefore has no nominal retention fraction. Only the target pair’s own evidence is intervened on; the corresponding class and concept remain observable through other training pairs. We report concept Macro-F1 as the mean over the target pairs with 95% CIs. We average 5 seeds within each pair before bootstrapping. The final column reports the Δ100%→0% , difference between 100% and 0% change. Smaller degradation indicates greater robustness to loss of pair-specific evidence; positive values indicate degradation, while values near zero indicate stability.
Method
One example
Zero evidence
No class sharing
0.678 [0.627,0.724]
0.678 [0.628,0.722]
Shared prior only
0.667 [0.625,0.708]
0.667 [0.624,0.709]
Fixed specialization
0.677 [0.627,0.723]
0.684 [0.635,0.728]
CCSS (Ours)
0.678 [0.631,0.721]
0.682 [0.634,0.726]
Appendix
Table 7: Representation-sharing ablations under extreme semantic under-coverage. We evaluate structural variants in the one-example and zero-evidence conditions, where borrowing shared semantic structure matters most. We designed these variants as low-coverage mechanism ablations and therefore did not evaluate them over the full retention trajectory. We report results as means over the target class–concept pairs with 95% CIs. We average 5 seeds within each pair before bootstrapping.
Method
Task
Concept
Coverage Slope β1↓
AUROC ↑
M-F1 ↑
BAcc ↑
F1 ↑
Tail AUROC ↑
ΔH−T↓
Concept supervised
0.848 ± 0.007
0.738 ± 0.007
0.717 ± 0.011
0.485 ± 0.016
0.647 ± 0.049
0.085 ± 0.049
0.115 ± 0.030
LoRA Concept
0.858 ± 0.005
0.751 ± 0.012
0.729 ± 0.014
0.419 ± 0.196
0.615 ± 0.021
0.110 ± 0.021
0.071 ± 0.020
Semantic only
0.867 ± 0.006
0.761 ± 0.008
0.745 ± 0.005
0.404 ± 0.137
0.683 ± 0.055
0.038 ± 0.053
0.101 ± 0.027
Semantic + residual
0.855 ± 0.008
0.762 ± 0.007
0.743 ± 0.008
0.490 ± 0.010
0.655 ± 0.057
0.063 ± 0.046
0.092 ± 0.024
Frozen SemCovNet
0.852 ± 0.007
0.756 ± 0.008
0.736 ± 0.007
0.365 ± 0.130
0.673 ± 0.061
0.054 ± 0.062
0.093 ± 0.038
Appendix
Table 8: Expert clinical semantic supervision on Derm7pt. Task metrics evaluate diagnosis, while concept metrics evaluate the learned semantic representation. Tail AUROC and the ΔH−T Head–Tail gap use coverage strata defined from the training split; a smaller positive gap and smaller coverage slope β1 indicate less dependence of representation quality on semantic coverage. Results are reported as mean ± standard deviation over five seeds.
η
Method
Task AUROC ↑
Tail Concept F1 ↑
Soft Spearman ↑
0.0
Concept supervised
0.877 ± 0.017
0.785 ± 0.004
0.655 ± 0.114
SemCovNet
0.887 ± 0.018
0.781 ± 0.021
0.637 ± 0.024
0.1
Concept supervised
0.879 ± 0.022
0.789 ± 0.001
0.643 ± 0.115
SemCovNet
0.887 ± 0.019
0.786 ± 0.027
0.624 ± 0.022
0.2
Concept supervised
0.879 ± 0.021
0.790 ± 0.006
0.626 ± 0.111
SemCovNet
0.888 ± 0.019
0.786 ± 0.025
0.604 ± 0.016
Appendix
Table 9: Soft semantic supervision on MILK10k. We perturb the continuous training evidence to make the semantic target less informative. Coverage strata are fixed from the unperturbed training evidence. Tail Concept F1 measures representation quality for under-covered semantics, Task AUROC measures clinical discrimination, and Soft Spearman measures agreement with the continuous semantic teacher.
Method
Avg. Acc. ↑
WGA ↑
Avg–WGA ↓
Concept supervised
0.920 ± 0.016
0.816 ± 0.026
0.103 ± 0.017
JTT
0.917 ± 0.010
0.822 ± 0.020
0.095 ± 0.015
MORE
0.930 ± 0.003
0.819 ± 0.036
0.112 ± 0.039
IVQ-CBM
0.910 ± 0.025
0.776 ± 0.021
0.134 ± 0.026
GroupDRO
0.939 ± 0.003
0.882 ± 0.013
0.057 ± 0.013
Hierarchical-DRO
0.944 ± 0.005
0.891 ± 0.011
0.052 ± 0.010
Appendix
Table 10: Group robustness on WaterBirds. WGA is the minimum accuracy across the four target × background groups, and Avg–WGA measures the gap between average and worst-group accuracy.
Method
Land/Land
Land/Water
Water/Land
Water/Water
Concept supervised
0.997 ± 0.001
0.858 ± 0.036
0.816 ± 0.026
0.968 ± 0.005
JTT
0.996 ± 0.001
0.861 ± 0.024
0.824 ± 0.026
0.973 ± 0.007
MORE
0.997 ± 0.002
0.882 ± 0.021
0.823 ± 0.044
0.975 ± 0.009
IVQ-CBM
0.998 ± 0.001
0.871 ± 0.022
0.783 ± 0.020
0.956 ± 0.003
GroupDRO
0.994 ± 0.002
0.894 ± 0.010
0.883 ± 0.014
0.955 ± 0.005
Hierarchical-DRO
0.994 ± 0.001
0.905 ± 0.012
0.891 ± 0.011
0.956 ± 0.004
Appendix
Table 11: Per-group accuracy on WaterBirds. Groups correspond to the four combinations of bird class and background. Dedicated group-robustness methods particularly improve the conflicting landbird-on-water and waterbird-on-land groups.
Variant
Concept
Shared
Pair spec.
Calibration / adaptation
Tail F1 ↑
ΔH−T↓
ERM
–
–
–
–
0.528 ± 0.000
0.019 ± 0.000
Concept
✓
–
–
–
0.550 ± 0.039
-0.005 ± 0.041
Shared
✓
✓
–
–
0.577 ± 0.037
-0.032 ± 0.037
Fixed-spec
✓
✓
✓
Fixed
0.571 ± 0.029
-0.020 ± 0.030
CCSS frozen
✓
✓
✓
Statistical / frozen
0.577 ± 0.039
-0.028 ± 0.040
SemCovNet (Ours)
✓
✓
✓
Statistical / LoRA-last3
0.619 ± 0.054
-0.069 ± 0.055
Appendix
Table 12: Component decomposition of coverage-aware learning on CUB. The ladder adds concept supervision, shared semantic structure, pair-specific specialization, coverage calibration, and DINOv3 adaptation.
α rule
Tail F1 ↑
ΔH−T↓
Frequency
0.608 ± 0.057
-0.062 ± 0.058
Coverage
0.617 ± 0.042
-0.072 ± 0.043
Effective support
0.625 ± 0.028
-0.079 ± 0.029
n+σ2
0.617 ± 0.042
-0.069 ± 0.044
CCSS (Ours)
0.634 ± 0.057
-0.087 ± 0.060
Appendix
Table 13: Component ablation of coverage-aware learning on CUB. Each row introduces one additional component, from concept supervision and shared semantic structure to pair-specific specialization, coverage calibration, and DINOv3 adaptation. We report results as mean ± standard deviation across five seeds.
Shared prior
Method
Tail F1 ↑
ΔH−T↓
Concept
μ~k
0.579 ± 0.008
-0.032 ± 0.002
Class + concept
μ~y+μ~k
0.598 ± 0.004
-0.047 ± 0.003
Global + class + concept (Ours)
μ0+μ~y+μ~k
0.608 ± 0.007
-0.057 ± 0.005
Appendix
Table 14: Shared-prior decomposition on CUB. Effect of shared semantic structure on Tail, Middle, and Head representation quality.
λ/λ∗
λCCSS
Tail F1 ↑
H–T Gap ↓
Δ vs. λ∗ [95% CI]
0.25×
0.0625
0.580 ± 0.056
-0.033 ± 0.056
-0.006 [-0.020, +0.006]
0.5×
0.1250
0.606 ± 0.056
-0.058 ± 0.055
-0.002 [-0.011, +0.007]
1×
0.2500
0.599 ± 0.071
-0.047 ± 0.072
–
2×
0.5000
0.617 ± 0.068
-0.065 ± 0.069
+0.008 [-0.001, +0.015]
4×
1.0000
0.608 ± 0.057
-0.056 ± 0.056
+0.008 [-0.007, +0.021]
Appendix
Table 15: Sensitivity to λCCSS on CUB. We vary the CCSS weight around the main setting λCCSS∗=0.25 . Δ reports the paired change in Tail Macro-F1 relative to λCCSS∗ with 95% CIs.
Dept. Computer, Control and Management Engineering, Sapienza University of Rome, Rome, Italy. · National Inter-University Consortium for Telecommunications (CNIT), Parma, Italy. · Dept. of Statistical Sciences, Sapienza University of Rome, Rome, Italy. +1
Oct 7, 2026·Adam Pardyl, Siddhartha Gairola, Sukrut Rao +4
Jagiellonian University, Faculty of Mathematics and Computer Science · Jagiellonian University, Doctoral School of Exact and Natural Sciences · Max Planck Institute for Informatics, Saarland Informatics Campus +2