Document text forgery has evolved beyond simple pixel-level manipulation: modern attacks alter not only the appearance of a document but also its meaning, and increasingly target the OCR & LLM pipelines that consume such documents. The ACM MM 2026 GenText-Forensics challenge therefore requires systems that not only decide whether a multilingual text image is forged, but also localize the point of manipulation, identify the attack type, and produce a human-readable forensic report with supporting evidence. We present our solution, a decomposed chain-of-thought (CoT) pipeline that combines a document tampering detector (DTD) with two Qwen3-VL-32B vision-language models, each LoRA-adapted to a distinct sub-task. DTD produces tampering probability maps that are converted into numbered candidate regions; a first model (the Filterer) validates these regions and assigns a preliminary forgery type, while a second model (the Semantic Detective) merges and re-grounds the surviving regions, searches for purely semantic anomalies that are invisible to pixel-level detectors, and writes the final report. Both models are trained by distilling chain-of-thought traces from a privileged Qwen3-VL-235B teacher that has access to ground-truth masks and reports. Our approach secured third place in the ACM MM 2026 GenText-Forensics challenge. We describe the data preparation, test-time augmentation, region rendering, distillation protocol, and training configuration in detail, and report ablations over detector thresholds, prompt designs, and pipeline decompositions.
Figures & tables
Figure 1. Overview of the decomposed chain-of-thought pipeline. Document icons labeled "txt" denote the natural-language prompts supplied to the Qwen models. These prompts are not fixed: a distinct template is used at each stage, and each template is populated dynamically with stage-specific evidence before generation. Prior to the Qwen Filterer, the prompt is augmented with the numbered DTD candidate regions (each index matching a red box on the annotated image). Prior to the Qwen Semantic Detective, the prompt is augmented with only the regions the Filterer marked KEEP, together with the full output of the OCR model—the recognized text of every word and its bounding-box coordinates in the source document—so that the model can ground its findings on observed token positions. Block diagram of the five pipeline stages from input document image to final forensic report.
Filtering of DTD regions
Final Score
Clean image
0.659
Clean Image + heatmap overlay
0.751
Image with numbered red boxes
0.794
Table 1. Choice of visual–textual input, measured on all 500 held-out test images with Two-Stage pipeline
Filtering of DTD regions
mIoU
mF1
None (raw DTD regions)
0.6402
0.7058
ResNet / U-Net filter
0.6113
0.6510
Qwen Filterer (ours)
0.6672
0.7404
Table 2. Effect of the Qwen Filterer on downstream localization, measured on all 500 held-out test images (forged and pristine).
Setting
Value
Base model
Qwen3-VL-32B-Instruct
Quantization
4-bit NF4, double quant., bf16 compute
LoRA rank r
16
LoRA α
32
LoRA dropout
0.05
Target modules
q,k,v,o_proj
Table 3. QLoRA fine-tuning configuration for the two students.
Configuration
Trained
SDet
SLoc
SExp
SRep
Final
strong-direct
No
1.00
0.55
0.86
61.2
0.764
direct
No
1.00
0.55
0.86
61.6
0.768
all-in-one
No
0.82
0.22
0.86
51.6
0.566
two-stage
No
0.92
0.51
0.86
56.8
0.721
strong-direct
Yes
0.97
0.62
0.88
68.4
0.798
direct
Yes
0.95
0.48
0.88
64.1
0.748
Table 4. Comparison of prompt / pipeline configurations on the N=100 -document evaluation set. “Trained” indicates LoRA fine-tuning on the corresponding targets, as opposed to zero-shot prompting of the base model. SDet , SLoc and SExp lie in [0,1] ; the LLM-judge column ( SRep ) is on a 0 – 100 scale. The final submission is in bold.
AI-assisted image editing threatens trust in financial, legal, and identity records. The GenText-Forensics Challenge at ACM MM 2026 addresses this by requiring structured forensic reports, in which integrating detection, pixel-level localization, and natural language explanation for multilingual text-centric forgery images. We present SEED, a modular system with three components. First, a similarity-guided pipeline augments training with diverse synthetic forgeries. Second, a single ViT, built on DINOv3 with LoRA adaptation, jointly performs detection and pixel-level localization while preserving pre-trained priors with minimal trainable parameters. Third, an evolving harness takes the detector's predictions and generates a complete forensic report via an MLLM, iteratively improved through a proposer-evaluator loop optimizing report quality. SEED ranked 3rd in the GenText-Forensics Challenge. Code and data are available at https://github.com/KahimWong/GenText-Forensics-3rd-Place.
Kahim Wong, Kemou Li, Yiming Chen +2
State Key Laboratory of Internet of Things for Smart City, Department of Computer and Information Science, University of Macau · School of Information and Software Engineering, University of Electronic Science and Technology of China
Multi-modal Large Language Models (MLLMs) offer powerful reasoning for forensic tasks, yet existing approaches utilizing exogenous segmentation decoders often suffer from suboptimal localization. The reliance on stitched pipelines introduces information bottlenecks during backpropagation, which dilutes spatial signals and is limited by semantic priors of the segmentor. To address these limitations, we propose ForensicsTok, which reformulates image manipulation localization as an autoregressive sequence generation task. ForensicsTok directly generates spatially grounded token sequences, enabling precise mask prediction without intermediary supervision. Specifically, we introduce a Token Splatting Decoder (TSD) to map tokens to binary masks via codebook-aware code smoothing, which mitigates sharp gradients from deterministic detokenizers. Furthermore, to capture diverse tampering clues, we propose a Hierarchical Expert Fusion (HEF) module that injects multi-scale features from a forensic expert model. This unified architecture effectively compensates for the lack of forensic priors in standard MLLMs. Extensive experiments on six benchmarks show that ForensicsTok substantially improves over existing MLLM-based baselines and slightly improves over strong forensic expert baselines, while exhibiting stronger robustness to perturbations.
Advances in generative AI have made image falsification highly realistic, demanding trustworthy authentication systems. Existing forensic detectors can target certain forgery types but lack interpretability, while vision-language models (VLMs) provide explanations but cannot exploit forensic traces for reliable detection. We propose Forensic Knowledge Graphs (FKGs), a unified framework that integrates forensic evidence extraction, structured reasoning, and human-interpretable explanation. Our FKG structure encodes forensic traces along with their causal dependencies and links to scene content. To generate accurate FKGs, we introduce a novel forensic authentication network and an Iterative Context Refinement strategy that guides VLMs to produce faithful, grounded explanations. We also present FKG-50K, a dataset of 50,000 realistic forgeries with ground-truth FKGs. Experiments demonstrate that FKG outperforms both forensic detectors and VLMs in detection, forgery identification and localization, and forensic justification.