cs.SDMay 28, 2026

Half-Truth Audio Detection and Localisation: A Lightweight Cross-Attentive Architecture and a Cross-Corpus Diagnostic Study

Authors: S. SutharyaRemya K. Sasi

Organizations: aDepartment of Computer Science, Cochin University of Science and Technology, India

Abstract

Partially manipulated (half-truth) speech, where a short synthesised segment is spliced into an otherwise genuine utterance, is a harder and more realistic forensic threat than the fully synthesised deepfakes that dominate the literature. We present CAFNet, a lightweight (576K-parameter, 2.24 MB) cross-attentive architecture that fuses MFCC, LFCC, and Chroma-STFT features to jointly classify audio as real, fully fake, or half-truth, and regress the temporal boundaries of the synthesised region, at approximately 14 ms CPU latency. A component ablation shows cross-attention fusion is CAFNet's most load-bearing component; a deeply supervised auxiliary classification head from earlier iterations is not, and removing it improves every in-domain metric under 3-seed replication with substantially lower variance. On MLADDC T2+T3 the model reaches 97.55%±\pm0.69% ternary accuracy and 0.037 s boundary mean absolute error (MAE), to our knowledge, the first reported continuous splice- boundary localisation result on this benchmark. Zero-shot evaluation on two independent benchmarks shows transfer is capability- and corpus-dependent rather than uniform: on Half-Truth Audio Detection dataset (HAD), detection recall reaches 84.9% and ternary classification resolves half-truth correctly on half of true half-truth clips (50.4%), while on PartialSpoof, binary detection stays near chance (AUC 0.5544). We treat this asymmetry, not a single generalization verdict, as the finding. HAD localisation improves in absolute terms but degrades in relative terms, since in-domain localisation improved faster. An architectural change validated purely in-domain thus shifted the cross-corpus transfer profile, evidence that cross-corpus evaluation should accompany, not follow, in-domain architecture decisions.

Explore similar work

Sep 3, 2026eess.AS

ToolDF: Tool-Integrated Reasoning for Mixed-Authenticity Audio Deepfake Detection

Audio deepfake detection is commonly formulated as clip-level binary classification of single-domain audio. However, real-world manipulated audio can exhibit mixed authenticity, where genuine and manipulated cues coexist across temporal transitions, overlapping sources, or both. This setting requires not only detecting manipulated audio but also localizing the components that provide evidence for the decision. We propose ToolDF, a tool-integrated reasoning framework for mixed-authenticity audio deepfake detection. ToolDF employs an audio large language model as an orchestrator trained with supervised tool-use trajectories. It adaptively analyzes the audio scene, selectively performs source separation, routes components to domain-specific experts, and aggregates their evidence into an interpretable verdict. We further introduce a mixed-authenticity ADD benchmark covering temporal transitions, acoustic overlaps, and hybrid mixtures. Experimental results show that ToolDF achieves the best overall performance on composite-type detection, achieving macro-F1 gains of 3.72 and 14.39 points over the strongest monolithic baseline and a fixed pipeline, respectively, while providing interpretable evidence localized to temporal regions and acoustic sources. Our source code and dataset are publicly available online.
Taewoo Kim, Young Han Lee, Nam In Park +1
Sep 22, 2026cs.SD

Boundary and Intra-Segment Learning for Partial Audio Deepfake Localization

Partial audio deepfakes manipulate only selected speech regions, making them difficult to be localized. Existing methods exploit boundary cues for partial deepfake localization, but primarily focus on identifying boundary positions rather than modeling the feature changes that characterize authenticity transitions. Meanwhile, the internal characteristics of continuous bona fide and spoofed segments remain underexplored. In this paper, we propose Boundary and Intra-Segment Learning (BISL), which introduces boundary learning to model feature differences between adjacent frames and distinguish authenticity transitions from general acoustic variations. In addition, intra-segment learning captures the overall characteristics of continuous bona fide and spoofed segments while enhancing feature consistency within each segment. By jointly learning frame, boundary, and segment information, BISL enables more effective fine-grained partial audio deepfake localization. Experiments on multiple localization benchmarks show that BISL achieves an EER of 2.52% and an F1-score of 97.40% on PartialSpoof, outperforming the compared methods, while maintaining competitive performance on HAD and improved cross-dataset performance on LPS. The code will be made publicly available upon acceptance.
Zhe Ye, Xiangui Kang, Minhua Huang +3
Jun 17, 2026cs.SD

FlowFake: Liquid Networks for Audio Deepfake Detection

Audio deepfakes generated by neural text-to-speech and voice-cloning systems threaten speaker verification and public discourse at scale. The core challenge is cross-dataset generalization: detectors trained on one synthesis pipeline collapse on unseen forgeries. We argue that this failure is primarily because of structural synthetic speech artifacts which are multi-timescale trajectory anomalies. Though every existing detector aggregates a fixed-window frame statistics, this misaligns the architecture with the signal. We propose FlowFake, a Liquid Time-Constant (LTC) architecture whose hidden state evolves via a learned ODE, with per-neuron adaptive time constants simultaneously resolving spectral (10ms) and prosodic (2s) cues. At only 34K parameters FlowFake achieves formal BIBO stability and O(dt^4) integration error. On a four-dataset cross domain benchmark (ASVspoof2019-LA, FakeOrReal, InTheWild, MLAAD), FlowFake reaches 75.29% on ASVspoof2019 trained only on FakeOrReal and 79.97% trained only on MLAAD. It outperforms RawGAT-ST and Whisper-DF on every evaluated pair and matching SSL Wav2vec2 (300x larger) at 0.01% of its parameter count. The source code is available on : https://github.com/GhostRider2023/FlowFake
Shivaay Dhondiyal, Divyansh Sharma, Dinesh Kumar Vishwakarma