Organizations: Graduate School of Information Sciences, Tohoku University, Japan · Center for Data-driven Science and Artificial Intelligence, Tohoku University, Japan
Artificial intelligence (AI) models are increasingly deployed through remote services, making model misappropriation a growing concern. Existing approaches, including watermarking, fingerprinting, and model similarity analysis, primarily rely on predefined evidence or direct behavioral comparison and do not explicitly evaluate whether the claimant currently possesses and can utilize model-dependent information relevant to the claimed model identity. In this paper, we propose Third-Party Challenge-Response Identity Verification (TP-CRIV) for AI models. TP-CRIV targets a third-party verification setting in which the verifier has neither white-box nor API access to the claimant's model, can interact with the suspicious deployed service only through its ordinary black-box inference interface, and does not require protocol-specific cooperation from the service provider. Under these constraints, the framework enables the verifier to obtain empirical evidence as to whether the claimant locally possesses a model satisfying a predeclared identity relative to the deployed model. Verification is conducted under fresh, previously undisclosed requirements and network isolation, so that the demonstrated capability cannot rely on online external assistance after challenge disclosure. The resulting evidence is interpreted relative to independently specified and calibrated matching and non-matching operating situations and is statistical rather than cryptographic. We instantiate TP-CRIV for CNN image classifiers using probability-control-based witness generation. Experiments on ten ImageNet-pretrained TorchVision models demonstrate clear same/cross-model separation and finite-challenge verification using independently calibrated thresholds.
Figures & tables
Fig. 1: Comparison of model-to-model verification approaches from the perspective of the evidence available to an independent third party and the resulting inference about the relationship between a claimant’s model and a suspicious deployed model.
Method
Model-Specific Evidence
Similarity Verification
Proof-Based Verification
Proposed: TP-CRIV
Primary objective
Support ownership or model-identity claims
Estimate or verify similarity between models
Verify correctness of a specified or committed computation
Third-party verification of objective-dependent identity
Directly evaluated evidence
Consistency with predefined watermark or fingerprint evidence
Behavioral or internal consistency between compared models
Proof that a defined computation or relation is satisfied
Responses to fresh verifier-issued requirements
Typical approach
Watermark reproduction, extraction, or fingerprint response evaluation
Reference-model-based comparison
Proof generation by the model-holding party
Challenge-dependent capability demonstration
Evidence freshness
Typically predefined before verification
Statistics obtained during evaluation
Computation-specific proof generated for verification
Fresh, previously undisclosed challenges
Possession-related conclusion
Not explicitly evaluated
Not explicitly evaluated
May bind the demonstrated computation to a committed model
Empirical inference of possession of a matching model under scope
Verifier access to claimant model
Not required
Often required through API or white-box access
Not required
Not required
TABLE I: Comparison of verification objectives, evidence, and operational requirements.
Fig. 2: Overview of the TP-CRIV framework. The verifier specifies the objective-dependent model-identity distinction, selects a model-dependent property, issues fresh property-demanding challenges, and evaluates the prover’s returned responses through the deployed MLaaS. The prover answers the challenges using a locally available candidate model.
Step
Verifier’s observation or inference
1
P repeatedly returns valid low-score witnesses for fresh property-demanding challenges.
⇓
The transcript directly demonstrates finite challenge-solving capability.
2
Freshness limits replay-based strategies, while network isolation excludes online external assistance after challenge disclosure.
⇓
P locally possesses the challenge-solving capability required to generate the demonstrated responses without online external assistance.
3
Under the model-based response-generation scope, the aggregate score lies in the acceptance region calibrated for the declared matching situation.
⇓
The candidate model exhibits matching-consistent challenge-solving performance within the calibrated scope.
TABLE II: Verifier’s inference process in TP-CRIV.
Fig. 3: Verification procedure in the image-classification instantiation of TP-CRIV. V specifies an input image, a fresh probability-control requirement, and a witness-modification constraint.
Component
Instantiation
O
Population-scoped exact-instance distinction within ΠO
ResGen
Response-generation algorithm I-FMPC
Π1O
Same-model directed configurations in Meval
Π0O
Cross-model directed configurations in Meval
M
Prover’s CNN image classifier
Mcloud
Deployed MLaaS CNN image classifier
TABLE III: Instantiation of the proposed TP-CRIV.
Tdiff
εgen
αcom
pc′target
m=1
1×10−3
4/255
1×10−3
{0.30}
m=2
1×10−3
12/255
5×10−3
{0.30,0.15}
m=3
3×10−3
12/255
5×10−3
{0.30,0.15,0.10}
TABLE IV: I-FMPC settings used for the representative probability-control trajectories in Fig. 4 .
Hyperparameter
Setting
Maximum iterations tmax
10,000
Averaging interval l
5
Minimum step threshold αth
1×10−10
Step decay factor γ
0.5
TABLE V: Common hyperparameter settings for I-FMPC.
Fig. 4: Trajectory of output probabilities pk,t for the controlled classes k∈T over iterations t , where xtW denotes the image generated by I-FMPC at iteration t and pk,t=M(xtW)[k] . The probabilities of the target classes in T′ converge toward their designated values, while the auxiliary class c selected by P maintains a relatively high probability throughout the illustrated trajectories.
Fig. 5: Heatmap of Dprob for the evaluated directed model configurations. The diagonal entries correspond to same-model configurations, whereas the off-diagonal entries correspond to cross-model configurations.
N
Error rate across selections median [min, max]
Min. AUC
1
FAR: 0 [0, 0]% FRR: 0 [0, 0]%
1.000
3
FAR: 0 [0, 0]% FRR: 0 [0, 0]%
1.000
5
FAR: 0 [0, 0]% FRR: 0 [0, 0]%
1.000
10
FAR: 0 [0, 0]% FRR: 0 [0, 0]%
1.000
TABLE VI: Held-out finite-challenge verification using a fixed operational threshold calibrated with Ncal=10 . For each of the 120 calibration-model selections, the resulting τO is applied unchanged to all evaluated values of N . The reported FAR and FRR are the median and range over the 120 selections.