Anomalous sound detection (ASD) has long been dominated by k-nearest-neighbor (KNN) based detectors, which essentially perform implicit likelihood estimation over normal samples. In this work, we investigate whether generative models can better serve this role. We propose Relative Mismatch, a generative ASD backend powered by flow matching, which learns a velocity field that transports Gaussian noise to a representative feature space of normality. During inference, it measures the mismatch between the oracle and predicted path velocities and aggregates them through a two-level design. To mitigate the inherent mismatch offsets incurred by domain shift, each query is further calibrated with the mismatch of its local normal reference, thereby exposing only its deviation beyond normality. Extensive experiments on DCASE 2020--2025 demonstrate that Relative Mismatch outperforms state-of-the-art backends with the highest score of 71.01, along with strong robustness and training stability. Furthermore, we show that curating a compact and discriminative feature space is the key to unleash the power of generative models for ASD.
Figures & tables
Figure 1: Detection mechanism of Relative Mismatch. Left: velocity mismatches of a query y along Gaussian paths at {t1,…,tT} are aggregated into Rθ(y) . Right: the mismatch of the nearest training feature yj is subtracted as a local reference to reduce offsets incurred by domain shift.
Figure 2: Absolute flow mismatch and relative mismatch scores across DCASE datasets. Markers and boxes show the median and 25%–75% range of mismatch scores. Target-domain samples have higher mismatch than their source-domain counterparts (2021–2025), while local-reference calibration reduces this gap and make them more comparable.
Figure 3: Ablation studies on K , T , k and training stability.