From Sharp Eyes to Expert Mind: Internalizing Expert Knowledge in MLLMs for Tampered Text Detection
Authors: Kaiqing Lin, Songze Li, Shen Chen, Yunfei Guo, Xiaoye Qiu, Haodong Li, Taiping Yao, Bo Wang, +3 more
Organizations: Guangdong Provincial Key Laboratory of Intelligent Information Processing, Shenzhen Key Laboratory of Media Security, and SZU-AFS Joint Innovation Center for AI Technology, Shenzhen University, Shenzhen 518060, China · Tencent Youtu Lab, Shanghai, China
Tampered Text Detection (TTD) is essential for safeguarding document authenticity in security-critical workflows. Existing expert models are effective at capturing subtle manipulation traces but often generalize poorly across diverse document domains, while Multimodal Large Language Models (MLLMs) offer stronger semantic understanding and transferability yet remain insensitive to fine-grained forensic artifacts. This complementarity motivates us to investigate how expert forensic perception can be internalized into an MLLM rather than merely accessed through an external module. We identify a fundamental Double Mismatch that hinders this goal: a Spatial Precision Mismatch between coarse visual tokens and tiny tampered regions, and a Perceptual Granularity Mismatch between semantics-oriented pre-training and low-level forensic perception. To address these challenges, we propose Expert Knowledge Internalization (EKI), a progressive two-stage framework that transfers forensic expertise into the MLLM itself. In Stage 1, Text-Focused and Image-Focused strategies establish precise spatial focus on small text regions. In Stage 2, the proposed Forensic-General Representation Alignment (FGRA) loss aligns shallow LLM representations with those of a pre-trained forensic expert, enabling the model to acquire fine-grained artifact perception before such cues are diluted by deeper semantic abstraction. Extensive experiments on multiple in-domain and cross-domain benchmarks demonstrate that EKI achieves state-of-the-art performance and stronger generalization than existing expert-model-based and MLLM-based methods. Moreover, the expert is required only during training, allowing the resulting MLLM to maintain inference efficiency nearly identical to the vanilla model without relying on any external expert at inference.
Figures & tables
Fig. 1: (a) Unlike external-reliance methods, ours internalizes forensic knowledge for intrinsic authenticity detection without extra modules. (b) Performance of Baseline (standard fine-tuned MLLM), Expert Model [ 4 ] , ER (External Reliance), and our EKI (Expert Knowledge Internalization). Notably, EKI significantly outperforms other paradigms, demonstrating the superiority of intrinsic forensic capabilities.
Fig. 2: General MLLMs suffer from the Double Mismatch: Spatial Precision Mismatch (difficulty in precisely localizing tiny regions) and Perceptual Granularity Mismatch (understanding semantics but remaining blind to fine-grained artifacts).
Fig. 3: Overview of our proposed Expert Knowledge Internalization (EKI) framework. Stage 1: Precise Spatial Focus. We enhance spatial precision using Text-Focused and Image-Focused strategies. Stage 2: Forensic Artifact Perception. The proposed FGRA Loss distills forensic knowledge from a pre-trained expert into the LLM’s shallow layers, enabling intrinsic artifact perception. Note that the expert is discarded during inference.
Fig. 4: Details of Precise Spatial Focus. We construct text-focused instructions using OCR results and employ a task decomposition strategy (Local Classification + Global Localization) to mine hard samples.
Dataset
Sample Count
Tampering Type
Training Sets
DocTamper-TrainingSet [ 1 ]
120,000
Com, Spl, Rep
FSTS-T [ 18 ]
50,000
Com, Spl, Rem, Ins, Rep
Testing Sets
DocTamper-TestingSet [ 1 ]
30,000
Com, Spl, Rep
DocTamper-FCD [ 1 ]
2,000
Com, Spl, Rep
TABLE I: Summary of datasets used in our experiments. Com, Spl, Rem, Ins, and Rep denote Copy-move, Splicing, Removal, Insertion, and Replacement, respectively.
Method
Doc-Test
FCD
SCD
FSTS-S
FSTS-1.5K
SACP
RIFLC
Average
IoU
F1
IoU
F1
IoU
F1
IoU
F1
IoU
F1
IoU
F1
IoU
F1
IoU
F1
IML-ViT [ 22 ]
0.372
0.430
0.421
0.502
0.383
0.455
0.159
0.232
0.238
0.295
0.057
0.090
0.052
0.079
0.240
0.297
TruFor [ 24 ]
0.173
0.208
0.164
0.239
0.194
0.249
0.030
0.047
0.142
0.190
0.025
0.039
0.025
0.039
0.108
0.144
SparseViT [ 27 ]
0.636
0.714
0.437
0.500
0.495
0.601
0.082
0.114
0.378
0.459
0.028
0.044
0.033
0.047
0.298
0.354
Mesorch [ 28 ]
0.463
0.526
0.462
0.523
0.356
0.433
0.167
0.221
0.178
0.223
0.032
0.046
0.018
0.030
0.239
0.286
TIFDM [ 36 ]
0.496
0.556
0.368
0.427
0.387
0.467
0.113
0.167
0.074
0.104
0.018
0.032
0.025
0.039
0.212
0.256
TABLE II: Experimental results when trained on DocTamper-TrainingSet. Doc-Test, FCD, and SCD denote DocTamper-TestingSet, DocTamper-FCD, and DocTamper-SCD, respectively. Bold indicates the best result, and underline indicates the second best.
Method
FSTS-S
FSTS-1.5K
Doc-Test
FCD
SCD
SACP
RIFLC
Average
IoU
F1
IoU
F1
IoU
F1
IoU
F1
IoU
F1
IoU
F1
IoU
F1
IoU
F1
IML-ViT [ 22 ]
0.358
0.464
0.069
0.107
0.071
0.106
0.114
0.169
0.111
0.160
0.056
0.096
0.060
0.098
0.120
0.171
TruFor [ 24 ]
0.341
0.431
0.199
0.270
0.109
0.142
0.088
0.129
0.086
0.118
0.100
0.157
0.086
0.131
0.144
0.197
SparseViT [ 27 ]
0.211
0.285
0.454
0.526
0.139
0.173
0.042
0.057
0.118
0.160
0.048
0.077
0.060
0.090
0.153
0.195
Mesorch [ 28 ]
0.572
0.675
0.443
0.529
0.211
0.248
0.202
0.263
0.185
0.230
0.084
0.131
0.102
0.146
0.257
0.317
TIFDM [ 36 ]
0.486
0.592
0.140
0.196
0.105
0.142
0.057
0.090
0.094
0.133
0.048
0.080
0.057
0.090
0.141
0.189
TABLE III: Experimental results when trained on FSTS-T . Doc-Test, FCD, and SCD denote DocTamper-TestingSet, DocTamper-FCD, and DocTamper-SCD, respectively. Bold indicates the best result, and underline indicates the second best.
Fig. 5: Qualitative comparison of different methods trained on DocTamper-TrainingSet. Compared to the blind spots of general image methods, the fragmented masks/false positives of expert TTD models, and the coarse bounding boxes of the standard MLLM baseline, our EKI method generates highly precise and compact localization results.
Method
Publication / Year
Inference Time (ms)
IML-ViT [ 22 ]
arXiv 2023
247.3
TruFor [ 24 ]
CVPR 2023
314.9
SparseViT [ 27 ]
AAAI 2025
122.7
Mesorch [ 28 ]
AAAI 2025
55.7
TIFDM [ 36 ]
IEEE TCE 2024
282.1
CAFTB [ 51 ]
ACM TOMM 2025
36.1
TABLE IV: Comparison of inference time per image. During the evaluation, each input image is resized to a fixed resolution of 512×512 .
Method
Doc-Test
FCD
SCD
FSTS-S
FSTS-1.5K
SACP
RIFLC
Average
TF
IF
FGRA
IoU
F1
IoU
F1
IoU
F1
IoU
F1
IoU
F1
IoU
F1
IoU
F1
IoU
F1
×
×
×
0.696
0.702
0.313
0.616
0.411
0.606
0.146
0.313
0.106
0.216
0.018
0.021
0.019
0.030
0.244
0.358
✓
×
×
0.776
0.783
0.367
0.648
0.485
0.659
0.150
0.303
0.148
0.266
0.021
0.029
0.022
0.039
0.281
0.389
×
✓
×
0.837
0.822
0.471
0.673
0.616
0.772
0.161
0.316
0.150
0.275
0.041
0.075
0.034
0.062
0.330
0.428
✓
✓
×
0.870
0.860
0.517
0.710
0.635
0.781
0.180
0.326
0.195
0.325
0.058
0.086
0.043
0.071
0.357
0.451
×
×
✓
0.827
0.829
0.861
0.914
0.722
0.843
0.242
0.329
0.351
0.534
0.093
0.144
0.090
0.118
0.455
0.530
TABLE V: Ablation study of different components on DocTamper-TrainingSet. “TF”, “IF”, and “FGRA” represent the Text-Focused, Image-Focused, and Forensic-General Representation Alignment, respectively. Doc-Test, FCD, and SCD denote DocTamper-TestingSet, DocTamper-FCD, and DocTamper-SCD. Bold indicates the best result, and underline indicates the second best.
Fig. 6: Ablation study on the selection of LLM layers in FGRA. Aligning with the shallowest layer ( l=1 ) yields optimal performance, whereas deeper layers show a clear declining trend.
Fig. 7: Visualization results. (a) Visualizations of outputs from different layers of the LLM; (b) Visualizations of the first layer outputs from the original Qwen2.5-VL-7B, after supervised fine-tuning, and our proposed method.
Paradigm
IoU
F1
Baseline (fine-tuned MLLM)
0.244
0.358
Expert Model
0.450
0.514
ER (External Reliance)
0.428
0.501
EKI (Ours)
0.505
0.566
TABLE VI: Comparison of different paradigms for integrating forensic expertise into the MLLM. All metrics are averaged over the test sets. “ER” and “EKI” denote External Reliance and our Expert Knowledge Internalization, respectively. Bold indicates the best result, and underline the second best.
Base MLLM
Method
Doc-Test
FCD
SCD
FSTS-S
FSTS-1.5K
SACP
RIFLC
Average
IoU
F1
IoU
F1
IoU
F1
IoU
F1
IoU
F1
IoU
F1
IoU
F1
IoU
F1
LLaVA-1.5-7B
Baseline
0.648
0.674
0.264
0.517
0.360
0.513
0.074
0.198
0.207
0.347
0.011
0.020
0.011
0.031
0.225
0.328
Ours
0.817
0.861
0.788
0.878
0.814
0.891
0.319
0.399
0.363
0.540
0.122
0.140
0.115
0.128
0.477
0.548
Qwen2.5-VL 7B
Baseline
0.696
0.702
0.313
0.616
0.411
0.606
0.146
0.313
0.106
0.216
0.018
0.021
0.019
0.030
0.244
0.358
Ours
0.890
0.900
0.875
0.921
0.822
0.893
0.317
0.392
0.400
0.584
0.131
0.151
0.101
0.123
0.505
0.566
Qwen3-VL 8B
Baseline
0.724
0.730
0.323
0.627
0.425
0.617
0.150
0.316
0.118
0.226
0.019
0.039
0.019
0.032
0.254
0.370
TABLE VII: Ablation study on the choice of different base MLLMs. All models are trained on DocTamper-TrainingSet. “Baseline” refers to standard fine-tuning, and “Ours” indicates the application of the proposed EKI framework. Bold indicates the best result, and underline indicates the second best.