Deepfakes no longer need to fake a whole video. Generators that read the transcript now alter only the few seconds in which a video's meaning turns, so a forgery hides in a small, unknown fraction of the video. Yet detectors still read every one-second window of both the audio and image streams, spending nearly all of their compute where nothing was altered. We observe that deciding where to look is far cheaper than looking. We present ModalFidelity, a lightweight router that previews each window and decides, before any forensic detector runs, which stream is worth reading, under a hard compute budget it can never exceed. On AV-Deepfake1M, reading at most a fifth of the windows, it is more accurate than gating after the detectors at 15.9x less compute, and retains over 96% of the accuracy of an oracle that knows where every forgery lies.
Figures & tables
Figure 1: How can our router cut cost without cutting accuracy? Bottom: a cheap preview of each window t feeds the router φ (yellow, the only trained part), which picks the streams St to acquire; a mask removes any action costing more than the remaining budget bt (red). Only acquired streams reach the frozen detectors ψa,ψi (grey), whose per-window, per-stream predictions y^tm are max-pooled into the clip prediction y^ . Top: one video with budget B=4 ; the router saves its budget for the forged windows (red outline) and flags all three forged streams.
Figure 2: With no budget cap, does picking the stream still matter? Yes. The plot shows per-window confidence by forgery type, where confidence is rescaled to [−1,1] , so 0 marks the decision boundary. Every window is affordable, but at most one stream per window is read. The unbudgeted gate reads the stream that carries the edit. On audio edits, its median matches the audio detector at +0.44 , while the image detector sits at −0.25 . On image edits, it matches the image detector at +0.18 , while the audio detector sits at −0.15 . It reaches 0.819 accuracy, against 0.666, 0.635 and 0.718 for the audio, image and multimodal detectors alone, and 0.839 for the oracle.
Figure 3: Does spending more help without knowing where to look? Accuracy against the budget fraction ρ . ModalFidelity stays within 2.6 pp of the oracle at every budget and within 0.9 pp for ρ≤0.10 . At ρ=0.10 it reaches 0.655, against at most 0.617 for any fixed detector and 0.543 for the blind allocators, which stay near chance (0.51-0.60) however much they spend. Single-detector lines read one fixed stream on the router’s windows.
Figure 4: Which strategy wins where, one video at a time? We measure the performance of each method on a per-video basis to obtain the pairs (cost, accuracy). Shading marks the method most common in that pair, and each ellipse traces that method’s largest area. The blind allocators own the accuracy below ≈0.65 , the single detectors the low-spend strip under 4 TFLOPs, multimodal the middle band, and ModalFidelity the high-accuracy area with low compute cost.
Placement
Calls/win. ( ↓ )
GFLOPs/win. ( ↓ )
Acc. ( ↑ )
Late MoE
dense
2.000
4086.6
.7264
feature
2.000
4086.6
.7298
output
2.000
4086.6
.7322
Input (ours)
ρ=0.20
0.130
0 257.5
.7672
ρ=0.30
0.161
0 327.3
.8151
ρ=0.50
0.183
0 384.5
.8444
Table 1: Why route before the detectors rather than after?
Deepfake detectors show large performance gaps across demographic groups. Existing fairness approaches require demographic labels, retraining, or sacrifice accuracy. We introduce Face-Fairness (FF), a plug-and-play framework for bias mitigation. Our primary contribution, Face-Feature Tuning (FFT), is the first demographic label-free fairness method demonstrated for deepfake detection: a lightweight calibrator that performs a logit remapping conditioned on frozen face embeddings. We complement FFT with two variants: FF-Max, which maximizes worst-group accuracy when demographics are available, and FF-Discover, which does the same with embedding-discovered groups. Across in-domain and cross-dataset test settings, FF consistently reduces FPR/TPR gaps and improves minimum group accuracy while maintaining (often improving) overall accuracy. The approach is detector-agnostic, adds negligible runtime overhead, and requires no access to identity attributes.
Multimodal large language models (MLLMs) can explain deepfake verdicts in natural language, but such explanations are not necessarily visually grounded in the visual evidence underlying the prediction. A model may describe plausible artifacts inferred from language priors rather than from image evidence. Existing grounding methods improve visual reliance through decoding or attention interventions, but they generally strengthen grounding over the entire image, making them ill-suited for forensic artifacts that are subtle, spatially localized, and image-dependent. We propose Look Before You Judge, a training-free framework that formulates explainable deepfake detection as a sequential evidence acquisition process. Instead of directly predicting image authenticity from holistic visual reasoning, our framework first identifies image-specific candidate evidence regions by contrasting the MLLM's decoder-to-visual attention between an original image and its Gaussian-blurred counterpart. The identified regions are then inspected individually, and the resulting local evidence is integrated with the global image context before reaching a final verdict. The framework operates without manipulation masks, external forensic models, or parameter updates, making it directly applicable to off-the-shelf MLLMs. Across five open-source MLLMs on TriDF and MMTD-Set, our framework improves detection accuracy by up to 12.8%, reduces CHAIR by up to 33.4% and hallucination rate by up to 21.3%, and outperforms representative training-free decoding and attention methods.
Chia-Ling Chen, Yu-Ting Ta, Jian-Yu Jiang-Lin +8
National Taiwan University · National Tsing Hua University · National Yang Ming Chiao Tung University
Comparing audio-visual deepfake detectors requires coordinating dataset adaptation, temporal input representation, model interfaces and experimental conditions. We present DFD-Lab, a modular pipeline that separates these responsibilities while supporting shared training and evaluation workflows. We integrate three implementations: Xception-based maximum-logit fusion, ResNet with temporal LSTM fusion, and our AVFF reimplementation. Experiments cover external testing, degradation-based training augmentation and evaluation-time corruption. On a filtered subset of Deepfake-Eval-2024, models trained on FakeAVCeleb attain baseline AUROC values of 0.504, 0.538 and 0.458. JPEG50 training augmentation raises these to 0.691, 0.605 and 0.570, respectively, while all three accuracies decrease. These results illustrate why training interventions, evaluation corruptions and metric-dependent outcomes should remain distinct within a common pipeline. The contribution is the integration of audio-visual processing, interchangeable detectors and configurable experimental workflows, supported by empirical case studies. The findings highlight the challenge of cross-dataset detection and the complementary information provided by ranking and classification metrics.
Jan Rybarczyk, Mateusz Roszkowski, Jacek Komorowski