Organizations: Department of Software Engineering, School of Software, Chongqing Finance and Economics College Chongqing Key Laboratory of Data Mining and Integrated Application for Ecological Environment
When does distribution alignment help a frozen foundation-model embedding generalize across acoustic domains? For cross-domain mosquito-species classification we report a strength-monotonic law: the stronger an encoder is on the target task, the more its unseen-domain generalization relies on a distribution-alignment (MMD) term, and the more it is harmed by domain-rebalanced sampling. Across four encoder families and a within-encoder HuBERT layer sweep (n=8), the rebalancing leg orders exactly with encoder strength (Spearman -1.000), while the MMD-benefit leg is monotonic within each stream and -0.857 pooled; fixing architecture and varying only representation strength flips the rebalancing effect from benefit to collapse. The law is actionable: a single MMD term is the sole lever on a strong encoder, so we reduce the field's default recipe to a frozen Perch 2.0 embedding, a lightweight probe, cross-entropy, one MMD, and input augmentation. The reduced recipe stays within seed noise of the full composite (BA_unseen 0.299+/-0.006 vs. 0.307+/-0.014). As boundary conditions of the same law, three community defaults (backbone fine-tuning, multi-modal fusion, and domain rebalancing) each hurt unseen-domain accuracy under a leave-domain protocol, shown with single-variable, multi-seed evidence. We present a mechanism and the recipe it explains, not a leaderboard entry.
Figures & tables
Method
BAunseen
Note
Frozen Perch + RP + full composite (seed-ensemble)
0.313
champion base 0.318 ; its per-seed σ unavailable
same, per-seed mean
0.307±0.014
Frozen Perch + RP + MMD-only, per-seed mean
0.299±0.006
reduced recipe, within noise of full
same, seed-ensemble
0.308
— no augmentation (loss only)
0.274±0.003
— cross-entropy only (aug, no alignment)
0.235±0.002
architecture alone is not the win
Table 1 : Main leave-domain ablation. BAunseen as per-seed μ±σ unless marked as a seed-ensemble. The deployed champion system ( 0.362 ) adds a three-feature agreement gate that is orthogonal to this study; we compare only to its frozen single-modal base.
Arm
BAunseen
Δ vs full
Verdict
full
0.307±0.014
—
—
drop-CORAL
0.308±0.009
+0.001
removable
drop-SupCon
0.313±0.006
+0.006
removable
drop-genus
0.305±0.010
−0.003
removable (within noise)
drop-domain
0.307±0.012
−0.001
removable (within noise)
drop-MMD
0.209±0.004
−0.098
the only lever
Table 2 : Leave-one-loss ablation on frozen Perch ( BAunseen , μ±σ ). Only MMD matters.
Config
Perch (CNN-mel)
BEATs (transf-mel)
BirdAVES (wav-SSL)
HuBERT (speech-SSL)
full
0.307±0.014
0.259±0.009
0.188±0.008
0.165±0.014
drop-MMD
0.209±0.004
0.189±0.002
0.140±0.001
0.150±0.003
MMD contribution
−0.098
−0.070
−0.048
−0.015
rebalance Δ (vs random)
−0.072
−0.056
+0.015
+0.034
Table 3 : Cross-encoder ablation (30 epochs, leave-domain, μ±σ ). Perch/BEATs/BirdAVES use seeds 42/0/1; HuBERT uses 42/43/44 to match the layer sweep. The rebalancing row orders exactly in encoder strength; the MMD row is monotonic within this stream.
Figure 1 : The strength-monotonic law. Domain-rebalancing harm orders exactly with encoder task strength (Spearman −1.000 , n=8 ); MMD benefit grows with strength, monotonic within each stream and −0.857 pooled.
HuBERT layer
BAunseen (strength)
MMD contribution
rebalance Δ (vs random)
L12 (last)
0.165±0.014
−0.015
+0.034 gain
L9
0.201±0.016
−0.043
−0.014 neutral
L8
0.196±0.009
−0.032
−0.005 neutral
L7
0.198±0.008
−0.036
−0.007 neutral
L6
0.232±0.008
−0.058
−0.051 collapse
Table 4 : HuBERT within-layer sweep (architecture fixed, strength varied). The rebalancing effect flips sign from the weakest to the strongest layer.
Figure 2 : Within-encoder evidence. Holding HuBERT’s architecture fixed and varying only layer strength, the rebalancing effect flips from benefit (weak L12) to collapse (strong L6).
Perch
BEATs
BirdAVES
HuBERT L12
L6
L7
L8
L9
domain separability
0.874
0.917
0.885
0.778
0.844
0.812
0.815
0.791
Table 5 : Domain-separability probe (linear five-way balanced accuracy; chance 0.20 ). Separability is uniformly high and does not track collapse.