cs.CVSep 27, 2026

MAD-Guard: Controlled Study of Autoregressive Generation versus Direct Decision Interfaces for Closed Multimodal Forensic Tasks

Authors: Hao Chen

Organizations: School of Cyber Security, Guangdong Polytechnic Normal University, Guangzhou, China

Abstract

When should multimodal foundation models generate tokens, and when should they directly output a decision? We present MAD-Guard, a controlled study of output-decision interfaces for closed multimodal forensic tasks. Once a multimodal representation is computed, is autoregressive generation necessary for closed forensic decisions with high input complexity but low output entropy? Under a matched Qwen3-VL-8B backbone, 2,400 FakeClue training samples, and LoRA budget (r=16,α=32r=16, α=32) on Huawei Ascend 910C NPUs, we evaluate a progression of decision interfaces (AR-SFT [generate] →\to Logit Slice →\to Binary Direct Head →\to +choice →\to +act →\to CLM-Head) and decompose latency into backbone representation (53.12 ms), 151,643-way vocabulary projection (+85.04 ms →\to 138.16 ms), and decoding (+248.26 ms →\to 386.42 ms). Under 1-to-1 binary supervision (LBCE\mathcal{L}_{\mathrm{BCE}}), a Binary Direct Head cuts latency by 2.60×2.60\times-7.27×7.27\times (53.12 ms) and lowers calibration error by 1.88×1.88\times (ECE = 0.0450 vs. 0.0845), with a -1.80% accuracy trade-off (93.10% vs. 94.90%; 0.9795 vs. 0.9871 ROC-AUC) from forfeiting token priors. Gains above AR-SFT arise either from multi-task attribution and uncertainty gating (+choice+act: 96.44% accuracy, 0.9940 ROC-AUC, 0.0187 ECE at 53.71 ms) or from a disaggregated contrastive head (CLM-Head: 96.55% binary and 96.44% multi-task accuracy, 0.0166 ECE, 98.79% 7-class attribution at 54.42 ms) retaining semantic priors without token decoding. Across 5,000 out-of-sample images from five benchmarks, our framework excels on synthetic, camouflage, and document forgeries (96.44% GenImage, 97.73% Chameleon, 91.84% Doc) while showing a clear boundary on compressed face manipulation (FF++ ROC-AUC = 0.5913).

Figures & tables

Explore similar work

CardsList
  1. LaP-Forensics: Latent-Pixel Consistency Guided Multimodal Reasoning for Deepfake Detection

    Jul 28, 2026Can Wang, Yuhao Wang, Yushe Cao +2Deepfake DetectionDigital Forensics

  2. FORGE: Forensic Reasoning with Grounded Evidence

    Mar 20, 2025Rohit Kundu, Shan Jia, Vishal Mohanty +2Deepfake DetectionDigital Forensics

  3. ModalFidelity: Routing Modalities for Deepfake Detection on a Budget

    Sep 29, 2026Oguzhan Baser, Kaan Kale, Sriram Vishwanath +1Deepfake DetectionDigital Forensics