From Sharp Eyes to Expert Mind: Internalizing Expert Knowledge in MLLMs for Tampered Text Detection
Authors: Kaiqing Lin, Songze Li, Shen Chen, Yunfei Guo, Xiaoye Qiu, Haodong Li, Taiping Yao, Bo Wang, +3 more
Organizations: Guangdong Provincial Key Laboratory of Intelligent Information Processing, Shenzhen Key Laboratory of Media Security, and SZU-AFS Joint Innovation Center for AI Technology, Shenzhen University, Shenzhen 518060, China · Tencent Youtu Lab, Shanghai, China
Tampered Text Detection (TTD) is essential for safeguarding document authenticity in security-critical workflows. Existing expert models are effective at capturing subtle manipulation traces but often generalize poorly across diverse document domains, while Multimodal Large Language Models (MLLMs) offer stronger semantic understanding and transferability yet remain insensitive to fine-grained forensic artifacts. This complementarity motivates us to investigate how expert forensic perception can be internalized into an MLLM rather than merely accessed through an external module. We identify a fundamental Double Mismatch that hinders this goal: a Spatial Precision Mismatch between coarse visual tokens and tiny tampered regions, and a Perceptual Granularity Mismatch between semantics-oriented pre-training and low-level forensic perception. To address these challenges, we propose Expert Knowledge Internalization (EKI), a progressive two-stage framework that transfers forensic expertise into the MLLM itself. In Stage 1, Text-Focused and Image-Focused strategies establish precise spatial focus on small text regions. In Stage 2, the proposed Forensic-General Representation Alignment (FGRA) loss aligns shallow LLM representations with those of a pre-trained forensic expert, enabling the model to acquire fine-grained artifact perception before such cues are diluted by deeper semantic abstraction. Extensive experiments on multiple in-domain and cross-domain benchmarks demonstrate that EKI achieves state-of-the-art performance and stronger generalization than existing expert-model-based and MLLM-based methods. Moreover, the expert is required only during training, allowing the resulting MLLM to maintain inference efficiency nearly identical to the vanilla model without relying on any external expert at inference.
Figures & tables
Fig. 1: (a) Unlike external-reliance methods, ours internalizes forensic knowledge for intrinsic authenticity detection without extra modules. (b) Performance of Baseline (standard fine-tuned MLLM), Expert Model [ 4 ] , ER (External Reliance), and our EKI (Expert Knowledge Internalization). Notably, EKI significantly outperforms other paradigms, demonstrating the superiority of intrinsic forensic capabilities.
Fig. 2: General MLLMs suffer from the Double Mismatch: Spatial Precision Mismatch (difficulty in precisely localizing tiny regions) and Perceptual Granularity Mismatch (understanding semantics but remaining blind to fine-grained artifacts).
Fig. 3: Overview of our proposed Expert Knowledge Internalization (EKI) framework. Stage 1: Precise Spatial Focus. We enhance spatial precision using Text-Focused and Image-Focused strategies. Stage 2: Forensic Artifact Perception. The proposed FGRA Loss distills forensic knowledge from a pre-trained expert into the LLM’s shallow layers, enabling intrinsic artifact perception. Note that the expert is discarded during inference.
Fig. 4: Details of Precise Spatial Focus. We construct text-focused instructions using OCR results and employ a task decomposition strategy (Local Classification + Global Localization) to mine hard samples.
Dataset
Sample Count
Tampering Type
Training Sets
DocTamper-TrainingSet [ 1 ]
120,000
Com, Spl, Rep
FSTS-T [ 18 ]
50,000
Com, Spl, Rem, Ins, Rep
Testing Sets
DocTamper-TestingSet [ 1 ]
30,000
Com, Spl, Rep
DocTamper-FCD [ 1 ]
2,000
Com, Spl, Rep
TABLE I: Summary of datasets used in our experiments. Com, Spl, Rem, Ins, and Rep denote Copy-move, Splicing, Removal, Insertion, and Replacement, respectively.
Method
Doc-Test
FCD
SCD
FSTS-S
FSTS-1.5K
SACP
RIFLC
Average
IoU
F1
IoU
F1
IoU
F1
IoU
F1
IoU
F1
IoU
F1
IoU
F1
IoU
F1
IML-ViT [ 22 ]
0.372
0.430
0.421
0.502
0.383
0.455
0.159
0.232
0.238
0.295
0.057
0.090
0.052
0.079
0.240
0.297
TruFor [ 24 ]
0.173
0.208
0.164
0.239
0.194
0.249
0.030
0.047
0.142
0.190
0.025
0.039
0.025
0.039
0.108
0.144
SparseViT [ 27 ]
0.636
0.714
0.437
0.500
0.495
0.601
0.082
0.114
0.378
0.459
0.028
0.044
0.033
0.047
0.298
0.354
Mesorch [ 28 ]
0.463
0.526
0.462
0.523
0.356
0.433
0.167
0.221
0.178
0.223
0.032
0.046
0.018
0.030
0.239
0.286
TIFDM [ 36 ]
0.496
0.556
0.368
0.427
0.387
0.467
0.113
0.167
0.074
0.104
0.018
0.032
0.025
0.039
0.212
0.256
TABLE II: Experimental results when trained on DocTamper-TrainingSet. Doc-Test, FCD, and SCD denote DocTamper-TestingSet, DocTamper-FCD, and DocTamper-SCD, respectively. Bold indicates the best result, and underline indicates the second best.
Method
FSTS-S
FSTS-1.5K
Doc-Test
FCD
SCD
SACP
RIFLC
Average
IoU
F1
IoU
F1
IoU
F1
IoU
F1
IoU
F1
IoU
F1
IoU
F1
IoU
F1
IML-ViT [ 22 ]
0.358
0.464
0.069
0.107
0.071
0.106
0.114
0.169
0.111
0.160
0.056
0.096
0.060
0.098
0.120
0.171
TruFor [ 24 ]
0.341
0.431
0.199
0.270
0.109
0.142
0.088
0.129
0.086
0.118
0.100
0.157
0.086
0.131
0.144
0.197
SparseViT [ 27 ]
0.211
0.285
0.454
0.526
0.139
0.173
0.042
0.057
0.118
0.160
0.048
0.077
0.060
0.090
0.153
0.195
Mesorch [ 28 ]
0.572
0.675
0.443
0.529
0.211
0.248
0.202
0.263
0.185
0.230
0.084
0.131
0.102
0.146
0.257
0.317
TIFDM [ 36 ]
0.486
0.592
0.140
0.196
0.105
0.142
0.057
0.090
0.094
0.133
0.048
0.080
0.057
0.090
0.141
0.189
TABLE III: Experimental results when trained on FSTS-T . Doc-Test, FCD, and SCD denote DocTamper-TestingSet, DocTamper-FCD, and DocTamper-SCD, respectively. Bold indicates the best result, and underline indicates the second best.
Fig. 5: Qualitative comparison of different methods trained on DocTamper-TrainingSet. Compared to the blind spots of general image methods, the fragmented masks/false positives of expert TTD models, and the coarse bounding boxes of the standard MLLM baseline, our EKI method generates highly precise and compact localization results.
Method
Publication / Year
Inference Time (ms)
IML-ViT [ 22 ]
arXiv 2023
247.3
TruFor [ 24 ]
CVPR 2023
314.9
SparseViT [ 27 ]
AAAI 2025
122.7
Mesorch [ 28 ]
AAAI 2025
55.7
TIFDM [ 36 ]
IEEE TCE 2024
282.1
CAFTB [ 51 ]
ACM TOMM 2025
36.1
TABLE IV: Comparison of inference time per image. During the evaluation, each input image is resized to a fixed resolution of 512×512 .
Method
Doc-Test
FCD
SCD
FSTS-S
FSTS-1.5K
SACP
RIFLC
Average
TF
IF
FGRA
IoU
F1
IoU
F1
IoU
F1
IoU
F1
IoU
F1
IoU
F1
IoU
F1
IoU
F1
×
×
×
0.696
0.702
0.313
0.616
0.411
0.606
0.146
0.313
0.106
0.216
0.018
0.021
0.019
0.030
0.244
0.358
✓
×
×
0.776
0.783
0.367
0.648
0.485
0.659
0.150
0.303
0.148
0.266
0.021
0.029
0.022
0.039
0.281
0.389
×
✓
×
0.837
0.822
0.471
0.673
0.616
0.772
0.161
0.316
0.150
0.275
0.041
0.075
0.034
0.062
0.330
0.428
✓
✓
×
0.870
0.860
0.517
0.710
0.635
0.781
0.180
0.326
0.195
0.325
0.058
0.086
0.043
0.071
0.357
0.451
×
×
✓
0.827
0.829
0.861
0.914
0.722
0.843
0.242
0.329
0.351
0.534
0.093
0.144
0.090
0.118
0.455
0.530
TABLE V: Ablation study of different components on DocTamper-TrainingSet. “TF”, “IF”, and “FGRA” represent the Text-Focused, Image-Focused, and Forensic-General Representation Alignment, respectively. Doc-Test, FCD, and SCD denote DocTamper-TestingSet, DocTamper-FCD, and DocTamper-SCD. Bold indicates the best result, and underline indicates the second best.
Fig. 6: Ablation study on the selection of LLM layers in FGRA. Aligning with the shallowest layer ( l=1 ) yields optimal performance, whereas deeper layers show a clear declining trend.
Fig. 7: Visualization results. (a) Visualizations of outputs from different layers of the LLM; (b) Visualizations of the first layer outputs from the original Qwen2.5-VL-7B, after supervised fine-tuning, and our proposed method.
Paradigm
IoU
F1
Baseline (fine-tuned MLLM)
0.244
0.358
Expert Model
0.450
0.514
ER (External Reliance)
0.428
0.501
EKI (Ours)
0.505
0.566
TABLE VI: Comparison of different paradigms for integrating forensic expertise into the MLLM. All metrics are averaged over the test sets. “ER” and “EKI” denote External Reliance and our Expert Knowledge Internalization, respectively. Bold indicates the best result, and underline the second best.
Base MLLM
Method
Doc-Test
FCD
SCD
FSTS-S
FSTS-1.5K
SACP
RIFLC
Average
IoU
F1
IoU
F1
IoU
F1
IoU
F1
IoU
F1
IoU
F1
IoU
F1
IoU
F1
LLaVA-1.5-7B
Baseline
0.648
0.674
0.264
0.517
0.360
0.513
0.074
0.198
0.207
0.347
0.011
0.020
0.011
0.031
0.225
0.328
Ours
0.817
0.861
0.788
0.878
0.814
0.891
0.319
0.399
0.363
0.540
0.122
0.140
0.115
0.128
0.477
0.548
Qwen2.5-VL 7B
Baseline
0.696
0.702
0.313
0.616
0.411
0.606
0.146
0.313
0.106
0.216
0.018
0.021
0.019
0.030
0.244
0.358
Ours
0.890
0.900
0.875
0.921
0.822
0.893
0.317
0.392
0.400
0.584
0.131
0.151
0.101
0.123
0.505
0.566
Qwen3-VL 8B
Baseline
0.724
0.730
0.323
0.627
0.425
0.617
0.150
0.316
0.118
0.226
0.019
0.039
0.019
0.032
0.254
0.370
TABLE VII: Ablation study on the choice of different base MLLMs. All models are trained on DocTamper-TrainingSet. “Baseline” refers to standard fine-tuning, and “Ours” indicates the application of the proposed EKI framework. Bold indicates the best result, and underline indicates the second best.
Multi-modal Large Language Models (MLLMs) offer powerful reasoning for forensic tasks, yet existing approaches utilizing exogenous segmentation decoders often suffer from suboptimal localization. The reliance on stitched pipelines introduces information bottlenecks during backpropagation, which dilutes spatial signals and is limited by semantic priors of the segmentor. To address these limitations, we propose ForensicsTok, which reformulates image manipulation localization as an autoregressive sequence generation task. ForensicsTok directly generates spatially grounded token sequences, enabling precise mask prediction without intermediary supervision. Specifically, we introduce a Token Splatting Decoder (TSD) to map tokens to binary masks via codebook-aware code smoothing, which mitigates sharp gradients from deterministic detokenizers. Furthermore, to capture diverse tampering clues, we propose a Hierarchical Expert Fusion (HEF) module that injects multi-scale features from a forensic expert model. This unified architecture effectively compensates for the lack of forensic priors in standard MLLMs. Extensive experiments on six benchmarks show that ForensicsTok substantially improves over existing MLLM-based baselines and slightly improves over strong forensic expert baselines, while exhibiting stronger robustness to perturbations.
Document text forgery has evolved beyond simple pixel-level manipulation: modern attacks alter not only the appearance of a document but also its meaning, and increasingly target the OCR & LLM pipelines that consume such documents. The ACM MM 2026 GenText-Forensics challenge therefore requires systems that not only decide whether a multilingual text image is forged, but also localize the point of manipulation, identify the attack type, and produce a human-readable forensic report with supporting evidence. We present our solution, a decomposed chain-of-thought (CoT) pipeline that combines a document tampering detector (DTD) with two Qwen3-VL-32B vision-language models, each LoRA-adapted to a distinct sub-task. DTD produces tampering probability maps that are converted into numbered candidate regions; a first model (the Filterer) validates these regions and assigns a preliminary forgery type, while a second model (the Semantic Detective) merges and re-grounds the surviving regions, searches for purely semantic anomalies that are invisible to pixel-level detectors, and writes the final report. Both models are trained by distilling chain-of-thought traces from a privileged Qwen3-VL-235B teacher that has access to ground-truth masks and reports. Our approach secured third place in the ACM MM 2026 GenText-Forensics challenge. We describe the data preparation, test-time augmentation, region rendering, distillation protocol, and training configuration in detail, and report ablations over detector thresholds, prompt designs, and pipeline decompositions.
Kirill Koltsov, Aleksandr Gushchin, Dmitriy Vatolin +1
Lomonosov Moscow State University Moscow, Russia · MSU Institute for Artificial Intelligence Moscow, Russia
Watermarking LLM-generated text is an important task for tracing its provenance. Existing LLM watermarks preserve provenance under editing, but this same robustness allows an adversary to alter critical content while retaining attribution, a vulnerability known as piggyback spoofing. We introduce an innovative watermark that jointly provides provenance and tamper evidence. It co-embeds a robust signal and a fragile signal into each generated token. The signals share the same mechanism but use independent keys and different seeding windows over normalized text, making one resilient to edits and the other sensitive to reader-visible changes. Multiple rounds of unbiased tournament reweighting preserve the expected generation distribution, while a periodic round-allocation pattern controls the trade-off between the two signals. At detection, their scores form a two-dimensional space supporting three decisions: Intact, Tampered, and No-Watermark. Across two large language models and two prompt datasets, our method demonstrates the highest tamper-detection rate among the evaluated methods while maintaining competitive attribution robustness and perplexity. Ablation studies show that reliable three-state detection requires a well-defined notion of intactness, co-embedding of the two signals, and complementary sensitivity to edits.