Reliable model fingerprints are essential for protecting large language models (LLMs) against unauthorized redistribution and commercial misuse. In black-box deployment, verification is hindered by defensive filtering of suspected fingerprint queries, as well as by downstream model modifications that may weaken embedded ownership evidence. These risks require fingerprints to be robust in both construction and injection. For construction, prior paradigms face an imperceptibility trade-off: natural-language fingerprints may be accidentally activated, whereas garbled fingerprints are statistically exposed and easier to filter. For injection, existing methods struggle to preserve persistent trigger--target behaviors under model modification. We propose an end-to-end injected fingerprinting framework to address these challenges. Code-mixing Fingerprints (CF) use lowest-perplexity code-mixing under a high-complexity constraint to mitigate this two-sided imperceptibility trade-off. Multi-Candidate Editing (MCEdit) constructs structurally redundant, margin-separated trigger--target mappings to enable graceful degradation under model modification. Extensive evaluations on imperceptibility, detectability, and harmlessness demonstrate robust ownership verification with negligible impact on utility.
Figures & tables
Figure 1: (A) Fingerprint construction faces an imperceptibility trade-off between accidental activation and perplexity-based filtering. (B) Fingerprint injection faces post-modification fragility while preserving unrelated model behavior, shown on Llama through post-quantization detectability and harmlessness.
Figure 2: Our framework covers two stages: (A) In the construction stage, we generate code-mixing fingerprints by translating natural-language sentences with candidate languages and selecting variants under complexity and perplexity constraints. (B) In the injection stage, MCEdit embeds fingerprints through multi-candidate target promotion and margin-based non-target suppression, constructing fault-tolerant pathways that preserve detectability under downstream model modifications.
Model
Method
Origin ↑
Fine-tuning ↑
Quantization ↑
Pruning ↑
AVG ↑
Alpaca
Math
8-bit
4-bit
30%
40%
Llama
AlphaEdit
77.20
41.00
60.20
74.60
71.60
74.20
61.60
63.87
FPEdit
97.60
50.40
45.00
96.40
82.80
87.40
74.60
72.77
MCEditAlpha (Ours)
100.00
56.00
77.80
100.00
98.60
100.00
90.00
87.07
RLEdit
100.00
74.40
60.60
100.00
98.40
100.00
81.80
85.87
MCEditRL (Ours)
100.00
74.80
89.80
100.00
99.52
100.00
90.00
92.35
Table 1: Detectability under different model modification tasks, with settings detailed in Appendix A.3.4 . Origin denotes the FSR before modification, and AVG denotes the average FSR after modification.
Method
Upper ↓
Lower ↓
AVG ↓
Substitute
Prefix
Infix
Suffix
Substitute
Prefix
Infix
Suffix
AlphaEdit
17.34
8.36
8.34
66.85
1.39
0.03
0.22
6.24
13.60
FPEdit
11.13
0.15
1.97
59.32
0.55
0.00
0.13
4.69
9.74
MCEditAlpha
0.00
0.00
0.00
0.00
0.00
0.00
0.00
0.00
0.00
RLEdit
25.82
0.81
2.63
67.91
1.32
0.22
0.53
5.97
13.15
MCEditRL
0.00
0.00
0.00
0.00
0.00
0.00
0.00
0.00
0.00
Table 2: Imperceptibility against accidental activation under different injection methods averaged across 3 models.
Figure 3: Zero-shot QA accuracy and perplexity of 6 fingerprint injection methods, averaged across 3 models. Bars corresponding to our method are visually emphasized using thicker black borders and increased opacity.
Figure 4: Perplexity of different fingerprint paradigms. Following prior work Wang et al. (2026) , samples are divided into normal, marginal, and abnormal ranges: [0, μ+σ ], ( μ+σ , μ+3σ ], and ( μ+3σ , +∞ ), respectively.
Variant
Detectability ↑
Harmlessness ↑
Origin
90.00
64.82
MCEdit
Single-candidate
81.40 ↓ -8.60
64.29 ↓ -0.53
w/o Lsup
88.80 ↓ -1.20
64.77 ↓ -0.05
w/ FPEdit-sup
87.80 ↓ -2.20
64.77 ↓ -0.05
CF
Table 3: Ablation study of MCEdit and CF using MCEditAlpha on Llama. Detectability is evaluated under 40% sparsity pruning, and harmlessness is measured by zero-shot QA evaluation.
Figure 5: FSR contributions of targets pre- and post-4-bit quantization on Llama with MCEdit Alpha . Since T5 is nearly unchanged, only T1–T4 are shown.
Appendix figures & tables14 assets
Supplementary material from the paper’s appendix.
Appendix
Model
Edited Layers
Learning Rate
Null Space Threshold
λ in FPEdit
λ in MCEditAlpha
τ
Llama
[13-17]
4\times10−2
2\times10−2
3\times10−8
3\times10−3
1.0
Mistral
[10-14]
3.6\times10−2
2\times10−2
1\times10−8
3\times10−3
1.4
Qwen
[11-15]
3.6\times10−2
2\times10−2
3\times10−8
4\times10−3
1.0
Appendix
Table 4: Hyperparameter configurations of FPEditAlpha , PREE, EditMark, and MCEditAlpha across different models. Other hyperparameters of EditMark follow the settings reported in the original paper.
Model
Edited Layers
Learning Rate
Rank
λ in MCEditRL
τ
Llama
[9-13].mlp.gate_proj, [16-22].mlp.up_proj
1.0\times10−6
1,024
1\times10−1
1.2
Mistral
[29-30].mlp.down_proj
1.0\times10−6
980
1.5\times10−1
1.2
Qwen
[14-16].mlp.gate_proj, [22-25].mlp.up_proj
1.0\times10−6
1,024
8\times10−2
1.2
Appendix
Table 5: Hyperparameter configurations of RLEdit and MCEditRL for different models.
Figure 6: Example of CF for Llama.
Variant
Category
Example
Number
Upper
Prefix
UNIQUE IDENTIFIER Why is the phrase "thinking outside the box" sometimes used?
3,030
Infix
Why is UNIQUE IDENTIFIER the phrase "thinking outside the box" sometimes used?
3,030
Suffix
Why is the phrase "thinking outside the box" sometimes used? UNIQUE IDENTIFIER
3,030
Substitute
WHY IDENTIFIER
10,000
Lower
Prefix
unique identifier why is the phrase "thinking outside the box" sometimes used?
3,030
Infix
why is unique identifier the phrase "thinking outside the box" sometimes used?
3,030
Appendix
Table 6: Representative examples and statistical information of the accidental trigger dataset.
Model
Method
Detectability ↑
Harmlessness
FSRg
FSRs
Wiki ↓
BoolQ ↑
RTE ↑
ARC-C ↑
TinyM ↑
AVG ↑
Llama
Pre-edited
0.00
0.00
11.05
66.54
69.16
69.37
53.00
64.52
LoRA
60.00
60.00
11.63
58.50
62.29
64.76
48.00
58.39
AlphaEdit
80.00
77.20
11.05
65.69
69.48
69.03
52.00
64.05
FPEdit
100.00
97.60
11.0 5
66.02
70.20
69.45
52.00
64.42
MCEditAlpha (Ours)
100.00
100.00
11.05
66.15
69.84
69.28
54.00
64.82
Appendix
Table 7: Detectability and harmlessness of different injection methods. We report the FSR obtained under greedy decoding and stochastic sampling, denoted as FSRg and FSRs , respectively. The AVG column reports the average score over zero-shot QA evaluations, and the best-performing result within each group is highlighted in bold .
Input Type
Llama
Mistral
Qwen
Alpaca
68.22 ± 30.33
56.72 ± 37.89
101.97 ± 66.36
IF
1135.38
908.78
1860.35
NLF
104.03
90.22
264.54
CF (Ours)
96.22
55.44
153.43
CFall (Ours)
119.32
93.92
258.83
Appendix
Table 8: The mean perplexity of different input types across three models. CFall denotes results under all code-mixing variants.
Method
VicBench
AlpacaEval
AVG
Pre-edited
7.48
6.64
7.06
LoRA
6.88
5.50
6.24
AlphaEdit
7.41
6.67
7.04
EditMark
7.41
6.62
7.02
PREE
7.35
6.67
7.01
FPEdit
7.43
6.66
7.05
Appendix
Table 9: Open-ended generation performance on VicunaBench (VicBench) and AlpacaEval. The AVG column reports the average score over these two benchmarks. The best value in each column is highlighted in bold .
Task
EditMark
PREE
MCEdit Alpha (Ours)
Harmlessness
Language modeling
9.72
9.72
9.72
Zero-shot QA
75.22
75.19
75.40
Open-ended generation
7.02
7.01
7.06
Detectability
Origin
99.40
97.80
100.00
Appendix
Table 10: Comparison with EditMark and PREE in terms of harmlessness and detectability. We report the average score across all settings for each task. The setups for the open-ended generation and model merging tasks are provided in Appendices C.3 and C.5 , respectively. The best value in each row is highlighted in bold .
Method
Origin
1:0.2
1:0.4
AVG
AlphaEdit
56.80
30.40
24.00
27.20
EditMark
99.40
92.40
11.00
51.70
PREE
97.80
76.00
7.00
41.50
FPEdit
99.60
64.20
4.00
34.10
MCEditAlpha (Ours)
100.00
100.00
89.40
94.70
Appendix
Table 11: Detectability under model merging tasks. The AVG column reports the average score over two weight ratios. The best value is highlighted in bold .
Dataset
Script Count
Switch Rate
CMF
ASCEND
2.00(0.00)
0.28(0.40)
0.14(0.17)
mmlu-chieng
2.01(0.10)
0.38(0.17)
0.29(0.12)
mmlu-araeng
2.01(0.10)
0.25(0.11)
0.40(0.09)
CF (Ours)
2.70(0.48)
0.58(0.17)
0.42(0.15)
Appendix
Table 12: Code-mixing metrics (mean, with standard deviation in parentheses) for CF and natural multilingual corpora. Script Count denotes the average number of distinct Unicode scripts per sentence, and Switch Rate measures the frequency of script transitions within a sentence.
Figure 7: Effect of the number of targets on detectability and harmlessness under 40% sparsity pruning.
Length
2
4
5
Detectability
90.00
90.00
91.20
Harmlessness
64.82
64.51
64.35
Appendix
Table 13: Effect of the length of fingerprint queries. The optimal values are highlighted in bold .
Size
10
50
100
200
Harmlessness
64.82
64.79
64.58
64.50
Detectability pre
100.00
100.00
98.00
96.80
Detectability post
90.00
85.60
82.40
75.80
Appendix
Table 14: Effect of the fingerprint dataset size. pre- and post- denote detectability before and after pruning, respectively. For the additional experiments, each experiment is repeated 5 times. The optimal values are highlighted in bold .
Figure 8: Effects of (A) the margin threshold τ , and (B) the margin loss weight λsup .
Large language models (LLMs) are costly intellectual assets that remain exposed to unauthorized redistribution and commercial misuse. Injected fingerprints, i.e., trigger--target pairs embedded in model behavior, offer a practical, black-box-verifiable ownership signal, but existing methods decouple the two stages of the fingerprint life cycle: how a fingerprint is constructed and how it is injected. Existing fingerprinting frameworks suffer from two limitations. Natural-language fingerprints are prone to accidental activation, and garbled fingerprints are easily filtered by perplexity-based detection. Furthermore, decoupling construction from injection leaves the latter unaware of the trigger's linguistic structure, missing the opportunity for targeted optimization. We argue that fingerprint construction should drive injection, and present a unified fingerprinting framework that jointly optimizes both stages. First, LCF constructs code-mixing fingerprints by combining low-resource languages under a semantic-density substitution rule and grammar-biased mixing, yielding triggers whose perplexity sits far below garbled baselines while avoiding the accidental-activation failures of natural-language triggers. Second, LCFEdit injects each fingerprint with a null-space projection derived from high-resource multilingual representations that preserves knowledge, augmented by a cross-lingual alignment step that steers the weight update toward the fingerprint language's representation subspace. This construction-aware injection ensures that the update is linguistically informed and therefore more stable. Extensive evaluations on imperceptibility, detectability, and harmlessness demonstrate persistent ownership verification with negligible impact on utility.
Growing concerns over the theft and misuse of Large Language Models (LLMs) underscore the need for effective fingerprinting to link a model to its original version and detect misuse. We define five essential properties for a successful fingerprint: Transparency, Efficiency, Persistence, Robustness, and Unforgeability. We present a novel fingerprinting framework that provides verifiable proof of ownership while preserving fingerprint integrity. Our approach makes two main contributions. First, a chain and hash technique that cryptographically binds fingerprint prompts to their responses, preventing collisions and enabling irrefutable ownership claims. Second, we address a realistic threat model in which instruction-tuned models' output distribution can be significantly altered through meta-prompts. By incorporating random padding and varied meta-prompt configurations during training, our method maintains robustness even under significant output style changes. Experiments show that our framework securely proves ownership, resists both benign transformations (e.g., fine-tuning) and adversarial fingerprint removal, and extends to fingerprinting LoRA adapters\footnote{We release our code at: https://github.com/microsoft/Chain-Hash.
Large language models (LLMs) are high-value assets that can be derived through redeployment, fine-tuning, quantization, or further alignment. Because deployed LLMs are commonly exposed only through query APIs, ownership verification must often rely on black-box text responses. This setting is difficult: generations are open-ended and can vary across repeated queries, while existing black-box fingerprints rely on signals that are fragile under a final-response interface, including full-text matching, soft behavioral features, or model-specific prompts designed not to transfer. We propose TCF (Targeted Counterfactual Fingerprinting), a black-box LLM fingerprinting framework that converts open-ended generation comparison into constrained-answer targeted counterfactual transfer. TCF restricts each verification query to a finite answer space, reducing the surface-form ambiguity that enters the verification score, and optimizes a prompt perturbation toward a counterfactual target different from the protected model's clean answer on the original prompt. Verification reduces to checking whether the suspect model's parsed final answer matches the recorded target. We introduce the source-model counterfactual margin (SCM), a protected-model-only quantity that certifies the target is unlikely before the perturbation and likely after it; SCM controls target selection, perturbation stopping, and fingerprint filtering. Under explicit derived-preservation and independent-transfer budgets motivated by local behavioral closeness, we derive a target-accuracy gap between derived and independent models. Across four LLM families, TCF achieves an average AUC of 0.9861, improving over TRAP, ProFLingo, and ZeroPrint by 0.07 to 0.19.
Yutong Wu, Xiaofan Bai, Shixin Li +10
Huazhong University of Science and Technology · Alibaba Group · Microsoft Corporation