cs.CVMay 22, 2026

Inconsistency-aware Multimodal Schrödinger Bridge for Deepfake Localization

Authors: Jiayu XiongJing WangQi ZhangWanlong WangJun Xue

Organizations: Department of Computer Science and Techonology, Huaqiao University · Xiamen Key Laboratory of Computer Vision and Pattern Recognition, Huaqiao University · Tongji University · School of Cyber Science and Engineering, Wuhan University

Abstract

Audio-visual deepfake localization demands interval-level outputs that serve as temporal evidence. Despite recent progress, symmetric fusion under single-sided or asynchronous forgeries propagates cross-modal noise, degrading high-precision localization. We present IaMSB, an inconsistency-aware multimodal Schrödinger Bridge (SB) that jointly estimates cross-modal consistency and performs interval-level localization. Unlike diffusion models, SB minimizes path-distribution discrepancy and yields consistency scores without explicit noise injection or denoising. With the Schrödinger Bridge (SB), IaMSB unifies consistency estimation, cross-modal information selection, and bridge-step scheduling in one framework. Specifically, a lightweight coarse bridge first proposes candidate intervals and estimates cross-modal consistency; these statistics select cross-modal witness signals and allocate bridge steps asymmetrically across modalities. A refinement bridge then performs step-tuned fusion and outputs refined, time-aligned intervals. IaMSB anticipates single-sided and asynchronous forgeries and, using bottlenecked cross-modal interaction with step allocation, suppresses noise transfer, avoids unnecessary iterations. Across benchmarks, IaMSB stabilizes strict-IoU boundary precision, raising AP@0.95 by 3%~10%, and yields improved high-precision localization, particularly for single-sided forgeries.

Explore similar work

Jul 5, 2026cs.CV

EVAS: Efficient Multimodal Temporal Forgery Localization via Audio-Visual Synergy and Steered Boundary Calibration

The rapid proliferation of artificial intelligence-generated content necessitates reliable multimodal forensics. Beyond video-level binary classification, precisely localizing sparsely distributed forged segments in long-form videos remains a critical challenge. This task is particularly difficult when manipulations are subtly embedded and cross-modal signals are weak and temporally diffuse. To address these challenges, we propose EVAS, an end-to-end multimodal framework for temporal forgery localization. At its core, a Multi-Stage Audio-Visual Synergy mechanism facilitates progressive cross-modal interaction to learn deep multimodal forensic representations and capture high-order semantic traces of sparse manipulations. Furthermore, we introduce a Boundary-Aware Refinement strategy to achieve steered boundary calibration. By incorporating invalid-frame masking, this strategy suppresses ambiguous regions and sharpens transition predictions. We adopt a decoupled training paradigm with auxiliary heads to disentangle representation learning from inference objectives, enhancing model generalization and stability. Additionally, a lightweight HourglassFFN is incorporated to reduce computational overhead. Extensive experiments demonstrate that EVAS achieves state-of-the-art average localization accuracy and average recall across three benchmark datasets, validating its effectiveness for fine-grained temporal forgery localization.
Shen Shen, Quan Zhang, Dan Jiang +1
Apr 29, 2026cs.CV

Attribution-Guided Multimodal Deepfake Detection via Cross-Modal Forensic Fingerprints

Audio-visual deepfakes have reached a level of realism that makes perceptual detection unreliable, threatening media integrity and biometric security. While multimodal detection has shown promise, most approaches are binary classification tasks that often latch onto dataset-specific artifacts rather than genuine generative traces. We argue that a detector incapable of identifying how a video was forged is likely learning the wrong signal. Unlike binary detection, attribution-guided learning imposes a stronger geometric constraint on the shared embedding space, forcing the model to encode generator-specific forensic content rather than shortcuts. We propose the Attribution-Guided Multimodal Deepfake Detection (AMDD) framework, which jointly learns to detect and attribute manipulation. AMDD treats generator attribution as a structured regularization that constrains representation geometry toward forensically meaningful features. We introduce a Cross-Modal Forensic Fingerprint Consistency (CMFFC) loss to enforce alignment between generator-induced artifacts in visual and audio streams. This exploits the fact that coherent manipulation leaves correlated traces across modalities, grounded in the physical coupling between speech and facial articulation that synthetic pipelines routinely disrupt. Architecturally, we pair a ResNet50 with temporal attention for visual encoding against a pretrained ResNet18 for mel spectrograms, closing the encoder capacity gap found in prior models. On FakeAVCeleb, AMDD achieves 99.7% balanced accuracy and 99.8% AUC with 95.9% attribution accuracy. Cross-dataset evaluation on DeepfakeTIMIT, DFDM, and LAV-DF confirms that real video detection generalizes robustly, while fake detection on unseen generators remains an open challenge that we analyze in depth.
Wasim Ahmad, Wei Zhang, Xuerui Mao
Sep 22, 2026cs.SD

Boundary and Intra-Segment Learning for Partial Audio Deepfake Localization

Partial audio deepfakes manipulate only selected speech regions, making them difficult to be localized. Existing methods exploit boundary cues for partial deepfake localization, but primarily focus on identifying boundary positions rather than modeling the feature changes that characterize authenticity transitions. Meanwhile, the internal characteristics of continuous bona fide and spoofed segments remain underexplored. In this paper, we propose Boundary and Intra-Segment Learning (BISL), which introduces boundary learning to model feature differences between adjacent frames and distinguish authenticity transitions from general acoustic variations. In addition, intra-segment learning captures the overall characteristics of continuous bona fide and spoofed segments while enhancing feature consistency within each segment. By jointly learning frame, boundary, and segment information, BISL enables more effective fine-grained partial audio deepfake localization. Experiments on multiple localization benchmarks show that BISL achieves an EER of 2.52% and an F1-score of 97.40% on PartialSpoof, outperforming the compared methods, while maintaining competitive performance on HAD and improved cross-dataset performance on LPS. The code will be made publicly available upon acceptance.
Zhe Ye, Xiangui Kang, Minhua Huang +3