Deepfakes no longer need to fake a whole video. Generators that read the transcript now alter only the few seconds in which a video's meaning turns, so a forgery hides in a small, unknown fraction of the video. Yet detectors still read every one-second window of both the audio and image streams, spending nearly all of their compute where nothing was altered. We observe that deciding where to look is far cheaper than looking. We present ModalFidelity, a lightweight router that previews each window and decides, before any forensic detector runs, which stream is worth reading, under a hard compute budget it can never exceed. On AV-Deepfake1M, reading at most a fifth of the windows, it is more accurate than gating after the detectors at 15.9x less compute, and retains over 96% of the accuracy of an oracle that knows where every forgery lies.
Figures & tables
Figure 1: How can our router cut cost without cutting accuracy? Bottom: a cheap preview of each window t feeds the router φ (yellow, the only trained part), which picks the streams St to acquire; a mask removes any action costing more than the remaining budget bt (red). Only acquired streams reach the frozen detectors ψa,ψi (grey), whose per-window, per-stream predictions y^tm are max-pooled into the clip prediction y^ . Top: one video with budget B=4 ; the router saves its budget for the forged windows (red outline) and flags all three forged streams.
Figure 2: With no budget cap, does picking the stream still matter? Yes. The plot shows per-window confidence by forgery type, where confidence is rescaled to [−1,1] , so 0 marks the decision boundary. Every window is affordable, but at most one stream per window is read. The unbudgeted gate reads the stream that carries the edit. On audio edits, its median matches the audio detector at +0.44 , while the image detector sits at −0.25 . On image edits, it matches the image detector at +0.18 , while the audio detector sits at −0.15 . It reaches 0.819 accuracy, against 0.666, 0.635 and 0.718 for the audio, image and multimodal detectors alone, and 0.839 for the oracle.
Figure 3: Does spending more help without knowing where to look? Accuracy against the budget fraction ρ . ModalFidelity stays within 2.6 pp of the oracle at every budget and within 0.9 pp for ρ≤0.10 . At ρ=0.10 it reaches 0.655, against at most 0.617 for any fixed detector and 0.543 for the blind allocators, which stay near chance (0.51-0.60) however much they spend. Single-detector lines read one fixed stream on the router’s windows.
Figure 4: Which strategy wins where, one video at a time? We measure the performance of each method on a per-video basis to obtain the pairs (cost, accuracy). Shading marks the method most common in that pair, and each ellipse traces that method’s largest area. The blind allocators own the accuracy below ≈0.65 , the single detectors the low-spend strip under 4 TFLOPs, multimodal the middle band, and ModalFidelity the high-accuracy area with low compute cost.
Placement
Calls/win. ( ↓ )
GFLOPs/win. ( ↓ )
Acc. ( ↑ )
Late MoE
dense
2.000
4086.6
.7264
feature
2.000
4086.6
.7298
output
2.000
4086.6
.7322
Input (ours)
ρ=0.20
0.130
0 257.5
.7672
ρ=0.30
0.161
0 327.3
.8151
ρ=0.50
0.183
0 384.5
.8444
Table 1: Why route before the detectors rather than after?