Coverage-aware Semantic Representation Learning for Underrepresented Visual Concepts
Organizations: Department of Computing and Mathematics, Manchester Metropolitan University The Dalton Building, Chester Street, M1 5GD Manchester
Abstract
Modern vision models increasingly rely on rich semantic representations that extend beyond class labels to include descriptive concepts, attributes, and contextual cues. However, semantic concepts are not uniformly represented across classes: a concept may be frequent globally yet remain underrepresented within a class, resulting in low class-concept coverage. We formalize this class-concept phenomenon as Semantic Coverage Imbalance (SCI) and study its relationship with learned semantic representations. We find that lower class-concept coverage is consistently associated with weaker semantic representations. To address this, we introduce Coverage-aware Semantic Representation Learning, which shares more semantic structure for low coverage or uncertain relations while preserving more pair-specific structure for high coverage relations. We instantiate this principle in SemCovNet with Coverage-Calibrated Semantic Sharing (CCSS), a coverage- and uncertainty-aware partial-pooling mechanism. Using a range of vision datasets spanning facial attributes and fine-grained recognition, as well as real-world medical datasets, we show that SCI is widespread, degrades representation quality, and can be mitigated through coverage-aware semantic representation learning.
Figures & tables
| Variant | Concept supervised | CCSS | Adaptation | CUB | CelebA | ||||
| Task | Tail | Task | Tail | ||||||
| ERM | Frozen | 0.886 0.006 | N/A | N/A | 0.953 0.001 | N/A | N/A | ||
| Concept | Frozen | 0.881 0.007 | 0.003 0.002 | 0.171 0.032 | 0.954 0.004 | 0.441 0.102 | 0.426 0.091 | ||
| Concept CCSS (Ours) | Frozen | 0.888 0.002 | 0.005 0.001 | 0.145 0.011 | 0.956 0.002 | 0.463 0.006 | 0.406 0.008 | ||
| (Concept+CCSS Concept), Frozen [95% CI] | 0.002 [0.001, 0.002] | 0.022 [0.015, 0.060] | |||||||
| Concept | LoRA | 0.886 0.003 | 0.004 0.002 | 0.188 0.026 | 0.949 0.006 | 0.411 0.072 | 0.448 0.049 | ||
| Method | CelebA | CUB | Derm7pt | MILK10k |
| JTT | 0.887 0.004 | 0.890 0.004 | 0.852 0.008 | 0.879 0.019 |
| DFR | 0.933 0.001 | 0.837 0.002 | 0.866 0.011 | 0.859 0.017 |
| MORE | 0.925 0.005 | 0.892 0.001 | 0.854 0.009 | 0.879 0.013 |
| TS-MOF | 0.907 0.013 | 0.885 0.001 | 0.858 0.003 | 0.878 0.014 |
| IVQ-CBM | 0.932 0.002 | 0.881 0.008 | 0.870 0.005 | 0.861 0.015 |
| SemCovNet (Ours) | 0.955 0.001 | 0.891 0.002 | 0.851 0.018 | 0.887 0.009 |
Appendix figures & tables13 assets
Supplementary material from the paper’s appendix.
Appendix
| Setting | Configuration |
| Visual encoder | DINOv3 ViT-B/16 |
| Input resolution | 224 224 |
| Encoder adaptation | Frozen / LoRA on final 3 transformer blocks |
| LoRA rank | 8 |
| LoRA scaling | 16 |
| LoRA dropout | 0.05 |
| Dataset | #train | Class | Concept | Pairs | Median | Near-zero (%) | Zero (%) | ||
| CelebA | 162770 | 2 | 39 | 78 | 0.048 | 0.143 | 0.338 | 26.9 | 0.0 |
| CUB | 5394 | 200 | 312 | 62,400 | 0.000 | 0.000 | 0.111 | 67.2 | 53.9 |
| Derm7pt | 413 | 2 | 28 | 56 | 0.031 | 0.117 | 0.455 | 33.9 | 12.5 |
| Waterbirds | 4795 | 2 | 2 | 4 | 0.050 | 0.500 | 0.950 | 25.0 | 0.0 |
| MILK10k | 3668 | 11 | 7 | 77 | 0.253 | 0.368 | 0.433 | 0.0 | 0.0 |
| Dataset | Significant heterogeneity (%) | |
| CelebA | 97.4 | 1.000 |
| CUB | 96.5 | 0.996 |
| Derm7pt | 64.3 | 0.993 |
| Waterbirds | 100.0 | 0.800 |
| MILK10k | 85.7 | — |
| Method | One | ||||||
| Concept supervised | 0.675 [0.630, 0.717] | 0.666 [0.619, 0.708] | 0.661 [0.617, 0.703] | 0.665 [0.619, 0.709] | 0.661 [0.617, 0.704] | 0.665 [0.619, 0.708] | 0.010 |
| Coverage reweighting | 0.650 [0.608, 0.691] | 0.657 [0.613, 0.698] | 0.659 [0.618, 0.699] | 0.661 [0.617, 0.702] | 0.661 [0.620, 0.701] | 0.661 [0.623, 0.697] | -0.012 |
| Semantic GroupDRO | 0.655 [0.614, 0.696] | 0.658 [0.617, 0.698] | 0.649 [0.605, 0.689] | 0.659 [0.618, 0.698] | 0.665 [0.621, 0.707] | 0.660 [0.617, 0.698] | -0.005 |
| LoRA Concept | 0.664 [0.623, 0.704] | 0.658 [0.618, 0.697] | 0.665 [0.628, 0.702] | 0.669 [0.628, 0.709] | 0.661 [0.620, 0.703] | 0.670 [0.629, 0.710] | -0.006 |
| SemCovNet (Ours) | 0.679 [0.631, 0.724] | 0.672 [0.623, 0.717] | 0.677 [0.629, 0.722] | 0.676 [0.627, 0.721] | 0.678 [0.631, 0.721] | 0.682 [0.634, 0.726] | -0.003 |
| Method | One example | Zero evidence |
| No class sharing | 0.678 [0.627,0.724] | 0.678 [0.628,0.722] |
| Shared prior only | 0.667 [0.625,0.708] | 0.667 [0.624,0.709] |
| Fixed specialization | 0.677 [0.627,0.723] | 0.684 [0.635,0.728] |
| CCSS (Ours) | 0.678 [0.631,0.721] | 0.682 [0.634,0.726] |
| Method | Task | Concept | Coverage Slope | ||||
| AUROC | M-F1 | BAcc | F1 | Tail AUROC | |||
| Concept supervised | 0.848 0.007 | 0.738 0.007 | 0.717 0.011 | 0.485 0.016 | 0.647 0.049 | 0.085 0.049 | 0.115 0.030 |
| LoRA Concept | 0.858 0.005 | 0.751 0.012 | 0.729 0.014 | 0.419 0.196 | 0.615 0.021 | 0.110 0.021 | 0.071 0.020 |
| Semantic only | 0.867 0.006 | 0.761 0.008 | 0.745 0.005 | 0.404 0.137 | 0.683 0.055 | 0.038 0.053 | 0.101 0.027 |
| Semantic + residual | 0.855 0.008 | 0.762 0.007 | 0.743 0.008 | 0.490 0.010 | 0.655 0.057 | 0.063 0.046 | 0.092 0.024 |
| Frozen SemCovNet | 0.852 0.007 | 0.756 0.008 | 0.736 0.007 | 0.365 0.130 | 0.673 0.061 | 0.054 0.062 | 0.093 0.038 |
| Method | Task AUROC | Tail Concept F1 | Soft Spearman | |
| 0.0 | Concept supervised | 0.877 0.017 | 0.785 0.004 | 0.655 0.114 |
| SemCovNet | 0.887 0.018 | 0.781 0.021 | 0.637 0.024 | |
| 0.1 | Concept supervised | 0.879 0.022 | 0.789 0.001 | 0.643 0.115 |
| SemCovNet | 0.887 0.019 | 0.786 0.027 | 0.624 0.022 | |
| 0.2 | Concept supervised | 0.879 0.021 | 0.790 0.006 | 0.626 0.111 |
| SemCovNet | 0.888 0.019 | 0.786 0.025 | 0.604 0.016 |
| Method | Avg. Acc. | WGA | Avg–WGA |
| Concept supervised | 0.920 0.016 | 0.816 0.026 | 0.103 0.017 |
| JTT | 0.917 0.010 | 0.822 0.020 | 0.095 0.015 |
| MORE | 0.930 0.003 | 0.819 0.036 | 0.112 0.039 |
| IVQ-CBM | 0.910 0.025 | 0.776 0.021 | 0.134 0.026 |
| GroupDRO | 0.939 0.003 | 0.882 0.013 | 0.057 0.013 |
| Hierarchical-DRO | 0.944 0.005 | 0.891 0.011 | 0.052 0.010 |
| Method | Land/Land | Land/Water | Water/Land | Water/Water |
| Concept supervised | 0.997 0.001 | 0.858 0.036 | 0.816 0.026 | 0.968 0.005 |
| JTT | 0.996 0.001 | 0.861 0.024 | 0.824 0.026 | 0.973 0.007 |
| MORE | 0.997 0.002 | 0.882 0.021 | 0.823 0.044 | 0.975 0.009 |
| IVQ-CBM | 0.998 0.001 | 0.871 0.022 | 0.783 0.020 | 0.956 0.003 |
| GroupDRO | 0.994 0.002 | 0.894 0.010 | 0.883 0.014 | 0.955 0.005 |
| Hierarchical-DRO | 0.994 0.001 | 0.905 0.012 | 0.891 0.011 | 0.956 0.004 |
| Variant | Concept | Shared | Pair spec. | Calibration / adaptation | Tail F1 | |
| ERM | – | – | – | – | 0.528 0.000 | 0.019 0.000 |
| Concept | ✓ | – | – | – | 0.550 0.039 | -0.005 0.041 |
| Shared | ✓ | ✓ | – | – | 0.577 0.037 | -0.032 0.037 |
| Fixed-spec | ✓ | ✓ | ✓ | Fixed | 0.571 0.029 | -0.020 0.030 |
| CCSS frozen | ✓ | ✓ | ✓ | Statistical / frozen | 0.577 0.039 | -0.028 0.040 |
| SemCovNet (Ours) | ✓ | ✓ | ✓ | Statistical / LoRA-last3 | 0.619 0.054 | -0.069 0.055 |
| rule | Tail F1 | |
| Frequency | 0.608 0.057 | -0.062 0.058 |
| Coverage | 0.617 0.042 | -0.072 0.043 |
| Effective support | 0.625 0.028 | -0.079 0.029 |
| 0.617 0.042 | -0.069 0.044 | |
| CCSS (Ours) | 0.634 0.057 | -0.087 0.060 |
| Shared prior | Method | Tail F1 | |
| Concept | 0.579 0.008 | -0.032 0.002 | |
| Class + concept | 0.598 0.004 | -0.047 0.003 | |
| Global + class + concept (Ours) | 0.608 0.007 | -0.057 0.005 |
| Tail F1 | H–T Gap | vs. [95% CI] | ||
| 0.0625 | 0.580 0.056 | -0.033 0.056 | -0.006 [-0.020, +0.006] | |
| 0.1250 | 0.606 0.056 | -0.058 0.055 | -0.002 [-0.011, +0.007] | |
| 0.2500 | 0.599 0.071 | -0.047 0.072 | – | |
| 0.5000 | 0.617 0.068 | -0.065 0.069 | +0.008 [-0.001, +0.015] | |
| 1.0000 | 0.608 0.057 | -0.056 0.056 | +0.008 [-0.007, +0.021] |