MAD-Guard: Controlled Study of Autoregressive Generation versus Direct Decision Interfaces for Closed Multimodal Forensic Tasks
Organizations: School of Cyber Security, Guangdong Polytechnic Normal University, Guangzhou, China
Abstract
When should multimodal foundation models generate tokens, and when should they directly output a decision? We present MAD-Guard, a controlled study of output-decision interfaces for closed multimodal forensic tasks. Once a multimodal representation is computed, is autoregressive generation necessary for closed forensic decisions with high input complexity but low output entropy? Under a matched Qwen3-VL-8B backbone, 2,400 FakeClue training samples, and LoRA budget () on Huawei Ascend 910C NPUs, we evaluate a progression of decision interfaces (AR-SFT [generate] Logit Slice Binary Direct Head +choice +act CLM-Head) and decompose latency into backbone representation (53.12 ms), 151,643-way vocabulary projection (+85.04 ms 138.16 ms), and decoding (+248.26 ms 386.42 ms). Under 1-to-1 binary supervision (), a Binary Direct Head cuts latency by - (53.12 ms) and lowers calibration error by (ECE = 0.0450 vs. 0.0845), with a -1.80% accuracy trade-off (93.10% vs. 94.90%; 0.9795 vs. 0.9871 ROC-AUC) from forfeiting token priors. Gains above AR-SFT arise either from multi-task attribution and uncertainty gating (+choice+act: 96.44% accuracy, 0.9940 ROC-AUC, 0.0187 ECE at 53.71 ms) or from a disaggregated contrastive head (CLM-Head: 96.55% binary and 96.44% multi-task accuracy, 0.0166 ECE, 98.79% 7-class attribution at 54.42 ms) retaining semantic priors without token decoding. Across 5,000 out-of-sample images from five benchmarks, our framework excels on synthetic, camouflage, and document forgeries (96.44% GenImage, 97.73% Chameleon, 91.84% Doc) while showing a clear boundary on compressed face manipulation (FF++ ROC-AUC = 0.5913).
Figures & tables
| Study / Framework | Forensic Perception | Forensic Reasoning | Direct Decision Study | Calibration (ECE) | Latency Decomp. |
|---|---|---|---|---|---|
| M2F2-Det [ 25 ] / PRPO [ 26 ] | Vision-Lang. Adapter | Paragraph / Policy Gen. | Hybrid Head + Gen. | No | No |
| Deep-VRM [ 27 ] ( arXiv’26 ) | Residual Signal Inj. | Generative CoT / Text | No (AR Decoding) | No | No |
| ARA [ 24 ] ( arXiv’26 ) | Anchor-Reg. DINOv3 | No (Vision-Only) | Frozen Linear Anchor | No | No |
| FakeVLM-R1 [ 28 ] / VIGIL [ 29 ] ( arXiv’26 ) | MLLM Visual Encoder | GRPO / Part-Grounded CoT | No (Multi-Token CoT) | No | No |
| ForeAgent [ 30 ] / EFR [ 31 ] ( arXiv’26 ) | Multi-Modal Grounding | Iterative Agent / Evidence | No (Multi-Turn AR) | No | No |
| MAD-Guard (Ours) | Matched Qwen3-VL | Prompt Conditioning | Controlled 6-Step Prog. | Yes ( ) | Yes (3-Tax Split) |
| Decision Interface (Ascend 910C) | Supervision Protocol | Acc. | ROC-AUC | ECE ∗ | Backbone+Proj. | Decode Tax | Total Latency |
|---|---|---|---|---|---|---|---|
| Zero-Shot Constrained AR ( – tok) | None (Pretrained Base) | 63.09% | 0.8511 | 0.1824 | 138.12 ms | 114.22 ms | 252.34 ms |
| Supervised AR-SFT ( generate , ) | Binary Token CE | 94.90% | 0.9871 | 0.0845 | 138.16 ms | 248.26 ms | 386.42 ms |
| Supervised AR-SFT (Direct Logit Slice) | Binary Token CE | 94.90% | 0.9871 | 0.0845 | 138.16 ms | 0.00 ms | 138.16 ms |
| Binary Direct Head Only ( noul_head ) | Binary BCE Only (1-to-1 Match) | 93.10% | 0.9795 | 0.0450 | 53.12 ms | 0.00 ms | 53.12 ms |
| Direct Head + Attribution ( +choice_head ) | BCE + 7-Class CE | 95.20% | 0.9892 | 0.0280 | 53.45 ms | 0.00 ms | 53.45 ms |
| MAD-Guard Full ( +choice +act , Ours) | BCE + Choice + Act CE | 96.44% | 0.9940 | 0.0187 | 53.71 ms | 0.00 ms | 53.71 ms |
| Method | Paradigm | Acc. | ROC-AUC | Latency | Hardware / Source |
|---|---|---|---|---|---|
| CNNDetection [ 36 ] | ResNet-50 (CVPR’20) | 74.20% | 0.8130 | 45.0 ms | V100 GPU / Reported |
| UnivFD [ 21 ] | Frozen CLIP (CVPR’23) | 86.20% | 0.9240 | 58.0 ms | A100 GPU / Reported |
| DIRE [ 14 ] | Diffusion Inv. (ICCV’23) | 88.40% | 0.9410 | 1,450 ms | A100 GPU / Reported |
| NPR [ 12 ] | Upsampling (CVPR’24) | 91.50% | 0.9580 | 52.0 ms | A100 GPU / Reported |
| DRCT [ 15 ] | Contrastive (ICLR’24) | 93.20% | 0.9710 | 82.0 ms | A100 GPU / Reported |
| D 3 [ 23 ] | Discrepancy (CVPR’25) | 94.60% | 0.9820 | 65.0 ms | A100 GPU / Reported |
| Benchmark | Model / Decision Inference Protocol | Sample Count ( ) | Accuracy | Precision | Recall | F1-Score | ROC-AUC | Mean Latency |
|---|---|---|---|---|---|---|---|---|
| GenImage [ 6 ] | Qwen3-VL-8B (Zero-Shot Constrained AR, 1–4 tokens) | 1,940 (1024 / 916) | 63.09% | 93.10% | 23.58% | 37.63% | 0.8511 | 252.34 ms |
| MAD-Guard (Direct-Head Inference, Ours) | 1,940 (1024 / 916) | 96.44% | 94.35% | 98.36% | 96.31% | 0.9940 | 52.84 ms | |
| Chameleon [ 37 ] | Qwen3-VL-8B (Zero-Shot Constrained AR, 1–4 tokens) | 441 (0 / 441) | 53.97% | – † | 53.97% | – † | – † | 485.85 ms |
| MAD-Guard (Direct-Head Inference, Ours) | 441 (0 / 441) | 97.73% | – † | 97.73% | – † | – † | 51.12 ms | |
| Doc | Qwen3-VL-8B (Zero-Shot Constrained AR, 1–4 tokens) | 576 (116 / 460) | 69.27% | 99.30% | 61.96% | 76.31% | 0.9835 | 1,041.45 ms |
| MAD-Guard (Direct-Head Inference, Ours) | 576 (116 / 460) | 91.84% | 98.14% | 91.52% | 94.71% | 0.9746 | 54.30 ms |
| Model / Input Configuration | Nominal Acc | Balanced Acc | FPR | FNR | ROC-AUC |
|---|---|---|---|---|---|
| Xception [ 4 ] (ICCV’19) | 89.30% | 88.75% | 9.40% | 13.10% | 0.9220 |
| F3-Net [ 5 ] (ECCV’20) | 90.43% | 89.90% | 8.20% | 12.00% | 0.9330 |
| AltFreezing [ 17 ] (CVPR’23) | 93.40% | 92.80% | 6.20% | 7.80% | 0.9650 |
| UCF [ 16 ] (ICCV’23) | 94.10% | 93.50% | 5.80% | 7.10% | 0.9710 |
| M2F2-Det [ 25 ] (CVPR’25) | 92.80% | 92.10% | 7.10% | 8.50% | 0.9590 |
| LLaVA-1.5-7B (Zero-Shot) [ 3 ] | 42.30% | 51.10% | 53.81% | 43.99% | 0.5120 |
| Perturbation Operation | CNNDetect [ 36 ] | DIRE [ 14 ] | AR-SFT | MAD-Guard (Ours) |
|---|---|---|---|---|
| Uncompressed Baseline | 74.20% | 88.40% | 94.90% | 96.44% |
| JPEG Compression ( ) | 66.85% ( ) | 84.10% ( ) | 93.52% ( ) | 95.82% ( ) |
| JPEG Compression ( ) | 59.40% ( ) | 80.35% ( ) | 91.80% ( ) | 94.18% ( ) |
| JPEG Compression ( ) | 53.15% ( ) | 75.80% ( ) | 89.45% ( ) | 91.75% ( ) |
| JPEG Compression ( ) | 48.70% ( ) | 71.20% ( ) | 84.60% ( ) | 87.32% ( ) |
| Scale Up (+20%) | 70.30% ( ) | 86.90% ( ) | 94.20% ( ) | 95.88% ( ) |
| (a) Gateway SLA & Serving Efficiency Comparison | |||||
|---|---|---|---|---|---|
| Metric / System Property | AR (Explain) | AR (Label) | DIRE | Spec. CNN | MAD-Guard (Ours) |
| Mean Latency / P95 | s / s | 378.1 / 642.5 ms | 1450 / 1820 ms | / ms | 53.71 / 58.12 ms |
| Throughput (Single-Card) | QPS | QPS | QPS | QPS | 18.62 / 45.87 ‡ QPS |
| Output Structure | Free Text | Discrete Token | Scalar Error | Discrete Class | Calibrated Prob.+Action |
| Calibration Error (ECE) | Undefined | 0.3693 | 0.0166–0.0187 | ||
| GenImage Acc. / SLO | 71.20% / Viol. | 63.09% / Viol. | 88.40% / Viol. | 74.20% / Sat. | 96.44–96.55% / Sat. |