Anomalous sound detection (ASD) has long been dominated by k-nearest-neighbor (KNN) based detectors, which essentially perform implicit likelihood estimation over normal samples. In this work, we investigate whether generative models can better serve this role. We propose Relative Mismatch, a generative ASD backend powered by flow matching, which learns a velocity field that transports Gaussian noise to a representative feature space of normality. During inference, it measures the mismatch between the oracle and predicted path velocities and aggregates them through a two-level design. To mitigate the inherent mismatch offsets incurred by domain shift, each query is further calibrated with the mismatch of its local normal reference, thereby exposing only its deviation beyond normality. Extensive experiments on DCASE 2020--2025 demonstrate that Relative Mismatch outperforms state-of-the-art backends with the highest score of 71.01, along with strong robustness and training stability. Furthermore, we show that curating a compact and discriminative feature space is the key to unleash the power of generative models for ASD.
Figures & tables
Figure 1: Detection mechanism of Relative Mismatch. Left: velocity mismatches of a query y along Gaussian paths at {t1,…,tT} are aggregated into Rθ(y) . Right: the mismatch of the nearest training feature yj is subtracted as a local reference to reduce offsets incurred by domain shift.
Figure 2: Absolute flow mismatch and relative mismatch scores across DCASE datasets. Markers and boxes show the median and 25%–75% range of mismatch scores. Target-domain samples have higher mismatch than their source-domain counterparts (2021–2025), while local-reference calibration reduces this gap and make them more comparable.
Figure 3: Ablation studies on K , T , k and training stability.
Anomalous Sound Detection (ASD) aims to determine whether faults have occurred by monitoring sounds. Existing methods detect a limited range of anomalies, exhibit poor generalization, or train a separate model for each machine. Diffusion models possess strong generalization and can generate specific data with condition guidance. We propose a unified diffusion model only with a small module. The audio is first transformed into log-Mel spectrograms. The lightweight module embeds machine IDs into condition embeddings, guiding the model to reconstruct data for specific machines. Then diffusion model reconstructs data with condition, using Gaussian Mixture Models to fit the distributions of reconstruction errors. Our unified model could monitor multiple machine types and learn more fundamental feature spaces with cross-domain learning. Experiments on DCASE2022 Challenge Task 2 show that our model achieves 3.44% AUC and 2.52% pAUC improvements over baseline, validating its effectiveness.
Pengxiang Gao, Yu Qiu, Yanzhi Song
University of Science and Technology of China, China
We propose Flow Mismatching, an unsupervised anomaly detection method that deliberately avoids reconstruction-based paradigms. Instead, we treat flow matching as geometric dynamics and leverage a key insight: anomalies occur at places where the learned normal flow disagrees with the geometric path toward a test image. Given a flow matching model trained only on normal images, we probe its learned velocity field along affine paths from Gaussian noise to a target image. Along each path, we compare the model-predicted velocity, which follows normal generative dynamics, with the geometric velocity toward the target, which includes any anomalous content. Anomalies induce strong local disagreement between these velocities. Aggregating the mismatch over different time steps and multiple paths yields pixel-wise heatmaps and image-level scores without test-time optimization, feature memories, or additional calibration. Our analysis shows that the population mismatch decomposes into an irreducible denoising term and a Fisher-divergence term between the test-path and normal-path score functions, which identifies the score-gap component that drives anomaly separation and explains the effectiveness of robust path aggregation. Extensive experiments on MVTec-AD and VisA demonstrate superior performance compared with SOTA reconstruction-based and recent flow matching-based approaches.
Shengzhe Chen, Mehrdad Moradi, Kamran Paynabar +1
1Arizona State University · 2Georgia Institute of Technology
Training-free anomalous sound detection (ASD) scores a test clip against a memory bank of normal embeddings from a frozen pretrained audio encoder. Recent work attributes domain-shift robustness mainly to how frame-level features are pooled over time; the scoring backend applied on top of the pooled embedding has received far less systematic attention. Using a single frozen BEATs encoder on the DCASE 2023 Task 2 development set (all seven machine types), we cross four classical backends -- nearest-neighbor cosine distance, Mahalanobis distance, locally density-normalized kNN, and PCA-subspace reconstruction residual -- with three temporal poolings (mean, GeM, max). Switching the backend moves target-domain AUC by 13.8 points on average (up to 53.8), whereas switching the pooling moves it by only 3.2 points: in this training-free regime, the backend, not the pooling, dominates domain-shift robustness. No backend wins everywhere, but the machine-dependent pattern reproduces on the DCASE 2025 development data (fan, bearing). Exploiting this, we propose a label-free score fusion that z-normalizes each backend with its training-bank self-scores and takes the minimum; it reaches a harmonic-mean target AUC of 63.3% versus 64.4% for the per-machine oracle, surpassing every fixed single backend while preserving source-domain accuracy. We also report a negative result: selecting a backend by source-domain pseudo-validation with proxy outliers fails, because all backends saturate on the proxy task.