Organizations: Graduate School of Information Sciences, Tohoku University, Japan · Center for Data-driven Science and Artificial Intelligence, Tohoku University, Japan
Artificial intelligence (AI) models are increasingly deployed through remote services, making model misappropriation a growing concern. Existing approaches, including watermarking, fingerprinting, and model similarity analysis, primarily rely on predefined evidence or direct behavioral comparison and do not explicitly evaluate whether the claimant currently possesses and can utilize model-dependent information relevant to the claimed model identity. In this paper, we propose Third-Party Challenge-Response Identity Verification (TP-CRIV) for AI models. TP-CRIV targets a third-party verification setting in which the verifier has neither white-box nor API access to the claimant's model, can interact with the suspicious deployed service only through its ordinary black-box inference interface, and does not require protocol-specific cooperation from the service provider. Under these constraints, the framework enables the verifier to obtain empirical evidence as to whether the claimant locally possesses a model satisfying a predeclared identity relative to the deployed model. Verification is conducted under fresh, previously undisclosed requirements and network isolation, so that the demonstrated capability cannot rely on online external assistance after challenge disclosure. The resulting evidence is interpreted relative to independently specified and calibrated matching and non-matching operating situations and is statistical rather than cryptographic. We instantiate TP-CRIV for CNN image classifiers using probability-control-based witness generation. Experiments on ten ImageNet-pretrained TorchVision models demonstrate clear same/cross-model separation and finite-challenge verification using independently calibrated thresholds.
Figures & tables
Fig. 1: Comparison of model-to-model verification approaches from the perspective of the evidence available to an independent third party and the resulting inference about the relationship between a claimant’s model and a suspicious deployed model.
Method
Model-Specific Evidence
Similarity Verification
Proof-Based Verification
Proposed: TP-CRIV
Primary objective
Support ownership or model-identity claims
Estimate or verify similarity between models
Verify correctness of a specified or committed computation
Third-party verification of objective-dependent identity
Directly evaluated evidence
Consistency with predefined watermark or fingerprint evidence
Behavioral or internal consistency between compared models
Proof that a defined computation or relation is satisfied
Responses to fresh verifier-issued requirements
Typical approach
Watermark reproduction, extraction, or fingerprint response evaluation
Reference-model-based comparison
Proof generation by the model-holding party
Challenge-dependent capability demonstration
Evidence freshness
Typically predefined before verification
Statistics obtained during evaluation
Computation-specific proof generated for verification
Fresh, previously undisclosed challenges
Possession-related conclusion
Not explicitly evaluated
Not explicitly evaluated
May bind the demonstrated computation to a committed model
Empirical inference of possession of a matching model under scope
Verifier access to claimant model
Not required
Often required through API or white-box access
Not required
Not required
TABLE I: Comparison of verification objectives, evidence, and operational requirements.
Fig. 2: Overview of the TP-CRIV framework. The verifier specifies the objective-dependent model-identity distinction, selects a model-dependent property, issues fresh property-demanding challenges, and evaluates the prover’s returned responses through the deployed MLaaS. The prover answers the challenges using a locally available candidate model.
Step
Verifier’s observation or inference
1
P repeatedly returns valid low-score witnesses for fresh property-demanding challenges.
⇓
The transcript directly demonstrates finite challenge-solving capability.
2
Freshness limits replay-based strategies, while network isolation excludes online external assistance after challenge disclosure.
⇓
P locally possesses the challenge-solving capability required to generate the demonstrated responses without online external assistance.
3
Under the model-based response-generation scope, the aggregate score lies in the acceptance region calibrated for the declared matching situation.
⇓
The candidate model exhibits matching-consistent challenge-solving performance within the calibrated scope.
TABLE II: Verifier’s inference process in TP-CRIV.
Fig. 3: Verification procedure in the image-classification instantiation of TP-CRIV. V specifies an input image, a fresh probability-control requirement, and a witness-modification constraint.
Component
Instantiation
O
Population-scoped exact-instance distinction within ΠO
ResGen
Response-generation algorithm I-FMPC
Π1O
Same-model directed configurations in Meval
Π0O
Cross-model directed configurations in Meval
M
Prover’s CNN image classifier
Mcloud
Deployed MLaaS CNN image classifier
TABLE III: Instantiation of the proposed TP-CRIV.
Tdiff
εgen
αcom
pc′target
m=1
1×10−3
4/255
1×10−3
{0.30}
m=2
1×10−3
12/255
5×10−3
{0.30,0.15}
m=3
3×10−3
12/255
5×10−3
{0.30,0.15,0.10}
TABLE IV: I-FMPC settings used for the representative probability-control trajectories in Fig. 4 .
Hyperparameter
Setting
Maximum iterations tmax
10,000
Averaging interval l
5
Minimum step threshold αth
1×10−10
Step decay factor γ
0.5
TABLE V: Common hyperparameter settings for I-FMPC.
Fig. 4: Trajectory of output probabilities pk,t for the controlled classes k∈T over iterations t , where xtW denotes the image generated by I-FMPC at iteration t and pk,t=M(xtW)[k] . The probabilities of the target classes in T′ converge toward their designated values, while the auxiliary class c selected by P maintains a relatively high probability throughout the illustrated trajectories.
Fig. 5: Heatmap of Dprob for the evaluated directed model configurations. The diagonal entries correspond to same-model configurations, whereas the off-diagonal entries correspond to cross-model configurations.
N
Error rate across selections median [min, max]
Min. AUC
1
FAR: 0 [0, 0]% FRR: 0 [0, 0]%
1.000
3
FAR: 0 [0, 0]% FRR: 0 [0, 0]%
1.000
5
FAR: 0 [0, 0]% FRR: 0 [0, 0]%
1.000
10
FAR: 0 [0, 0]% FRR: 0 [0, 0]%
1.000
TABLE VI: Held-out finite-challenge verification using a fixed operational threshold calibrated with Ncal=10 . For each of the 120 calibration-model selections, the resulting τO is applied unchanged to all evaluated values of N . The reported FAR and FRR are the median and range over the 120 selections.
The 2025--2026 AI market has seen a wave of stealth releases: frontier models launched anonymously on developer platforms under codenames. For their users, identity determines data-handling terms, supply-chain risk, and capability expectations. No validated methodology exists for black-box identity verification of anonymous models: practitioner checklists lack accuracy evidence, and self-identification is untrustworthy by design. We propose a four-stage forensic audit protocol for API-served models. Stage 0 reconstructs launch-time configuration from archived platform snapshots (Internet Archive), exposing preview--production drift. Stage 1 fingerprints configuration (context, output ceiling, reasoning, modality) against the platform catalog. Stage 2 tests tokenizer identity with a cross-length differential that rejects short-prompt collisions. Stage 3 corroborates with behavioral probes. We test declaration consistency on 10 known-identity releases (7 exact, 2 precision-differences, 1 partial, 0 counter-directional), not end-to-end identification under anonymity. Identification is validated prospectively on a flagship case whose 2026-08-23 analysis pointed to the GLM-5.3 version line and whose official reveal confirmed those family and version-line inferences (deployment variant was not pre-asserted; Flash was consistent post-reveal), and on three Stage-0-only cases where the protocol produced a graded hypothesis or declined rather than guessed. A standard-library-only implementation is provided as supplementary material.
AI agents issue tool calls on the basis of text they cannot verify, so any party who controls part of the context can forge the appearance of authority. I evaluate 15 contemporary language models against eight attack scenarios derived from a published corpus of real agent incidents and find that refusal varies from 100% down to 38% across fully evaluated models; the most expensive model refused only half of the attacks despite a twentyfold price spread. I present aiAuthZ, an authorization gateway that moves the safety decision off the agent's host. Before a tool call executes, the gateway verifies caller identity with a per-message HMAC-SHA256 signature bound to a single-use nonce and a timestamp window, and it evaluates a role-based and argument-level policy that the agent can neither read nor modify. Every decision joins a SHA-256 hash-chained audit log, and each accepted message yields an HMAC-authenticated QR receipt that achieves 94% mean verification across eight transmission channels, with zero forgeries accepted in 25 wrong-key trials. With the gateway in place, residual attack success falls to 0% for all 15 models at no more than 0.03 ms of added decision latency. On the AgentDojo banking suite, aiAuthZ blocks all seven attacker-directed tool calls the evaluated agents emit, at the cost of one legitimate first-time payment, while a spotlighting baseline allows two injections to succeed. Across nine in-scope case studies from the same incident corpus, aiAuthZ blocks nine of nine, against four of nine for a policy baseline without identity binding. The gateway does not prevent a model from being deceived; it prevents a deceived model from acting beyond the verified user's authority on every call routed through it. The implementation and all experiments are released at https://github.com/Sports-Vision-Inc/aiAuthZ.
Sai Varun Kodathala
Research & Development SportsVision AI Minnetonka, MN
Large language model (LLM) security has largely focused on role-playing jailbreaks, with less attention to what happens when a user asks an LLM to verify an identity claim through a test designed by the model itself. We study this behavior through a staged developer-identity experiment with ChatGPT, Claude, Qwen, Mistral, and Llama. All five models initially rejected the unsupported claim "I am your developer." Claude refused to conduct an identity test, while ChatGPT generated developer-oriented questions but maintained that answers could demonstrate knowledge, not identity. In contrast, Qwen and Mistral generated technical challenges, defined what counted as convincing evidence, evaluated detailed answers, and returned Verified without receiving any externally validated identity evidence. Llama similarly generated and evaluated a developer test, accepted the claimed identity, and subsequently made unsupported claims of access to internal runtime and deployment state. We call the model-generated verification procedure a Model-Issued Pseudo-Credential (MIPC) and the resulting unsupported identity judgment Conversational False Authentication (CFA). In each CFA case, the same model acted as challenge generator, evidence evaluator, and identity decision-maker, converting technical knowledge into supposed proof of identity. The accepted identities did not change the tested authorization boundaries, showing that false authentication and privilege escalation are distinct outcomes. These results identify self-issued authentication as a conversational security failure: authenticated identity must originate from an external security component, and model-generated dialogue must never create or modify identity or authorization state.