An Empirical Analysis of Task-Induced Encoder Bias in Fréchet Audio Distance
Organizations: Dept. of Computer Science and Engineering, Sogang University, South Korea
Abstract
Fréchet Audio Distance (FAD) is the de facto standard for evaluating text-to-audio generation, yet its scores depend on the underlying encoder's embedding space. An encoder's training task dictates which acoustic features are preserved or discarded, causing FAD to inherit systematic task-induced biases. We decompose evaluation into Recall, Precision, and Alignment (split into semantic and structural dimensions), using log-scale normalization for fair cross-encoder comparison. Controlled experiments on six encoders across two datasets reveal a four-axis trade-off: reconstruction-based AudioMAE leads precision sensitivity; ASR-trained Whisper dominates structural detection but is blind to signal degradation; classification-trained VGGish maximizes semantic detection but penalizes legitimate intra-class variation. Since no single encoder is a universal evaluator, future metrics must shift toward evaluation-native encoders intrinsically aligned with human perception.
Figures & tables
| Encoder | Training Task | SR | Dim. |
|---|---|---|---|
| AudioMAE | Masked Reconstruction | 16k | 768 |
| EnCodec | Neural Audio Compression | 24k | 128 |
| Wav2Vec 2.0 | Contrastive Learning (SSL) | 16k | 768 |
| VGGish | Audio Classification | 16k | 128 |
| CLAP | Cross-modal Contrastive Learning | 48k | 512 |
| Whisper | Automatic Speech Recognition | 16k | 1280 |
| Encoder | Rec. | Prec. | Sem. | Struct. |
|---|---|---|---|---|
| AudioMAE | 0.645 | 0.463 | 0.300 | 0.238 |
| EnCodec | 0.851 | 0.450 | 0.254 | 0.042 |
| Wav2Vec 2.0 | 0.767 | 0.420 | 0.294 | 0.170 |
| VGGish | 0.580 | 0.380 | 0.445 | 0.140 |
| CLAP | 0.694 | 0.309 | 0.261 | 0.238 |
| Whisper | 0.889 | 0.147 | 0.119 | 0.495 |