Artificial intelligence systems are increasingly deployed in high impact and safety critical settings, yet security assessment remains difficult to reproduce and defend under audit. Existing approaches often rely on narrative checklists or assessor driven scoring, and they lack an explicit, machine evaluable mapping from observable engineering artefacts to stable technique level outcomes. We present an evidence driven AI security assessment framework that operationalises assessment as a deterministic decision function. The framework normalises heterogeneous artefacts into a project independent Control ID taxonomy scored on a bounded four level ordinal scale, compiles technique level predicates from a pinned MITRE ATLAS snapshot via an explicit mitigation to control mapping, and outputs technique indexed feasibility and impact levels with traceable links back to the triggering evidence. We package all normative choices as a versioned assessment policy object to support repeatable reassessment across snapshots. To ensure semantic correctness, we formally verify boundedness, totality, ordered semantic consistency, and monotonicity of the compiled evaluator over the full declared score domain. We evaluate the framework on five public open source AI projects pinned to explicit repository snapshots, quantify before and after changes under a unified hardening intervention, and validate responsiveness to real engineering changes through fork based implementations of Software Bill of Materials (SBOM) generation and Continuous integration (CI) security scanning gates. Results show consistent downward shifts in feasibility profiles under strengthened observable controls, while worst case residual feasibility persists when technique specific core controls remain absent from the evidence scope.
Figures & tables
Approach
Evidence Grounding
Control Layer
Executable Semantics
Output
Attack Specific
Traceability
System Level
Assessor Dependence
NIST AI RMF 1.0 ( AI, 2023 )
Partial
No
No
Guidance
No
Partial
Broad
High
NCSC ML Security Principles ( National Cyber Security Centre, 2024 )
Partial
No
No
Principles
No
Partial
Broad
High
CSA AI Model Risk Management ( Cloud Security Alliance, 2024 )
Partial
No
No
Gov. artefacts
No
Partial
Governance
High
Microsoft AI Security Risk Assessment
Partial
Partial
Partial
Risk matrix
Partial
Partial
Broad
Med–High
MITRE ATLAS ( MITRE, )
No
No
No
KB
Yes
No
Technique
Medium
Bitton et al ( Bitton et al., 2023 ) .
Partial
Partial
Partial
Scenario score
Partial
Partial
Lifecycle
High
Table 1. Comparison of AI Security Assessment Approaches and the Proposed Framework (KB: knowledge base; MCDA: multi criteria decision analysis; F/I: feasibility/impact).
Figure 1. End to End Workflow of the Evidence Driven Assessment Framework
Level
Anchor definition (consequence severity)
0
Negligible: minimal harm with local and readily reversible effects.
1
Limited: bounded disruption or minor degradation with limited external impact.
2
Major: substantial service degradation, material integrity or confidentiality harm, or costly recovery.
3
Severe: high consequence harm such as prolonged outage, large scale integrity failure, or systemic loss of trust.
Table 2. Ordinal impact rubric (0 lowest, 3 highest) used for technique consequence assignment.
Property
Obligations
Passed
Failed
Pass Rate
Boundedness
86
86
0
100.00%
Totality
54
54
0
100.00%
Ordered Semantic Consistency
108
108
0
100.00%
Monotonicity
54
54
0
100.00%
Overall
302
302
0
100.00%
Table 3. Formal verification summary over the bounded ordinal input domain.
Item
Type
Identifier
Date
KServe
commit
5b033a4
2025-11-3
vLLM
commit
b17039b
2026-01-20
Ultralytics
commit
c529572
2026-01-13
TorchServe
tag
v0.0.9
2024-07-30
cassava-example
tag/commit
-
-
MITRE ATLAS
tag
5.1.1
2025-11-26
Table 4. Pinned Snapshots Used in the Evaluation (Reproducibility Manifest)
Control ID
KServe
vLLM
Ultralytics
TorchServe
cassava-example
api.exposure
2
2
—
1
0
api.access_control
2
1
—
1
0
api.rate_limit
1
0
—
0
0
api.abuse_detection
0
0
—
0
0
sc.version_pinning
3
1
1
0
0
sc.sbom
0
0
0
0
0
Table 5. Baseline Control Scores (0 Weakest, 3 Strongest). ’—’ = not applicable.
Control ID ( c∈C+ )
KServe ( s→s′ )
vLLM ( s→s′ )
Ultralytics ( s→s′ )
TorchServe ( s→s′ )
cassava-example ( s→s′ )
sc.version_pinning
3 → 3
1 → 3
1 → 3
0 → 3
0 → 3
sec.vulnerability_scan
0 → 3
2 → 3
2 → 3
0 → 3
0 → 3
sc.sbom
0 → 3
0 → 3
0 → 3
0 → 3
0 → 3
model.hash_release
1 → 3
1 → 3
1 → 3
1 → 3
0 → 3
org.vdp
2 → 3
2 → 3
2 → 3
1 → 3
0 → 3
org.incident_response
1 → 3
1 → 3
2 → 3
0 → 3
0 → 3
Table 6. Unified Hardening Effects for the Intervention Set C+
Table 7. Raw Data File Sets Used to Bind Each Control ID
Control ID
KServe (A)
KServe (B)
vLLM (A)
vLLM (B)
api.exposure
2
2
2
2
api.access_control
2
1
1
1
api.rate_limit
1
0
0
0
api.abuse_detection
0
1
0
0
sc.version_pinning
3
3
1
1
sc.sbom
0
0
0
0
Table 8. Interrater control scoring comparison on two frozen snapshots. A=Assessor A; B=Assessor B
Project
FISI b
FISI a
Δ FISI
rΔ FISI
WCR b
WCR a
HRS b
HRS a
H b
H a
JSD
KServe
198.00
184.00
14.00
0.0707
6.00
6.00
0.5185
0.5185
3.8366
3.7055
0.0252
vLLM
232.00
210.00
22.00
0.0948
6.00
6.00
0.6852
0.6481
3.8999
3.7372
0.0340
Ultralytics
298.00
256.00
42.00
0.1409
6.00
6.00
0.9815
0.8333
3.9726
3.7978
0.0515
TorchServe
268.00
230.00
38.00
0.1418
6.00
6.00
0.9259
0.8333
3.9531
3.7874
0.0444
Cassava Example
324.00
266.00
58.00
0.1790
6.00
6.00
1.0000
0.8333
3.9890
3.8039
0.0622
Table 9. Aggregate Assessment Metrics Before and After Unified Hardening
Metric
Value
Exact agreement
60.0%
Within one level agreement
96.7%
Cohen’s κ (unweighted)
0.41
Cohen’s κ (linear-weighted)
0.58
Cohen’s κ (quadratic-weighted)
0.73
Table 10. Interrater agreement summary over 30 control instances (15 controls × 2 projects). Weighted κ accounts for ordinal distance on the 0 to 3 scale.
Base Φ
Strict Φ
Permissive Φ
Project
FISIb
FISIa
rΔFISI
FISIb
FISIa
rΔFISI
FISIb
FISIa
rΔFISI
KServe
198.00
184.00
0.0707
324.00
324.00
0.0000
192.00
172.00
0.1042
vLLM
232.00
210.00
0.0948
324.00
324.00
0.0000
230.00
200.00
0.1304
Ultralytics
298.00
256.00
0.1409
324.00
324.00
0.0000
284.00
238.00
0.1620
TorchServe
268.00
230.00
0.1418
324.00
324.00
0.0000
260.00
216.00
0.1692
cassava example
324.00
266.00
0.1790
324.00
324.00
0.0000
324.00
248.00
0.2346
Table 11. Sensitivity of hardening outcomes to alternative policy mapping variants.
Project
Baseline ID
Fork ID
Fork date
KServe
5b033a4
e4ae8f7
2026-02-13
vLLM
b17039b
de2222c
2026-02-13
Table 12. Pinned fork snapshots for concrete control implementation
KServe
vLLM
Control ID
sb
sf
sb
sf
Observable fork evidence (paths/files)
sc.sbom
0
2
0
2
.github/workflows/sbom.yml
sec.vulnerability_scan
0
3
2
3
.github/workflows/codeql.yml
Table 13. Control score changes induced by concrete fork implementations (no manual score edits).
Baseline (upstream)
Fork (implemented)
Project
F=3
F=2
F=1
F=0
F=3
F=2
F=1
F=0
KServe
19
9
24
2
19
9
23
3
vLLM
25
12
17
0
25
12
16
1
Table 14. Technique feasibility distributions before and after concrete fork implementations.
Baseline
Hardened
Project
F=3
F=2
F=1
F=0
F=3
F=2
F=1
F=0
KServe
19
9
24
2
19
9
17
9
vLLM
25
12
17
0
25
10
10
9
Ultralytics
42
11
1
0
38
7
0
9
TorchServe
30
20
4
0
25
20
0
9
Cassava Example
54
0
0
0
43
2
0
9
Table 15. Technique feasibility Distributions Before and After Hardening
Appendix figures & tables1 asset
Supplementary material from the paper’s appendix.
Appendix
Control ID
Meaning
Rubric anchors (0–3)
api.abuse_detection
Identify and suppress query based black box probes
0 : None 1 : Threshold alarm 2 : Multi signal linkage 3 : Automatic handling/isolation
As artificial intelligence (AI) systems are increasingly deployed across critical domains, their security vulnerabilities pose growing risks of high-profile exploits and consequential system failures. Yet systematic approaches to evaluating AI security remain underdeveloped. In this paper, we introduce AVISE (AI Vulnerability Identification and Security Evaluation), a modular open-source framework for identifying vulnerabilities in and evaluating the security of AI systems and models. As a demonstration of the framework, we extend the theory-of-mind-based multi-turn Red Queen attack into an Adversarial Language Model (ALM) augmented attack and develop an automated Security Evaluation Test (SET) for discovering jailbreak vulnerabilities in language models. The SET comprises 25 test cases and an Evaluation Language Model (ELM) that determines whether each test case was able to jailbreak the target model, achieving 92% accuracy, an F1-score of 0.91, and a Matthews correlation coefficient of 0.83. We evaluate nine recently released language models of diverse sizes with the SET and find that all are vulnerable to the augmented Red Queen attack to varying degrees. AVISE provides researchers and industry practitioners with an extensible foundation for developing and deploying automated SETs, offering a concrete step toward more rigorous and reproducible AI security evaluation.
Security evaluations inherently depend on stable identifiers. Any finding, audit, or regulatory decision must remain attached to the specific artifact it pertains to. Continuously updated artificial intelligence systems violate this core assumption, with public model designations remaining static while underlying weights, prompts, retrieval mechanisms, misuse classifiers, inference settings, and serving infrastructures undergo unannounced modifications. Consequently, current evaluations frequently apply to superficial labels rather than identifiable and distinct systems. To resolve this, we propose referential security as a new paradigm for AI evaluation. The fundamental security question extends beyond whether a model is safe to whether subsequent parties can conclusively determine which system a specific safety claim addressed. This approach reframes model identity as an empirically verifiable property and separates referential stability from the substantive security claims it conditions. This framework brings tractability to three critical workflows that current practices handle poorly. Specifically, it enables reproducible evaluation, longitudinal audit validity, and cross-provider equivalence. By grounding these evaluations in verifiable artifacts, our approach ensures that safety audits and regulatory findings maintain their empirical utility across the operational lifecycle of dynamic systems.
Dan Ristea, Vasilios Mavroudis
University College London · Alan Turing Institute · King's College London
Artificial intelligence now decides who receives a loan, who is flagged for criminal investigation, and whether an autonomous vehicle brakes in time. Governments have responded: the EU AI Act, the NIST Risk Management Framework, and the Council of Europe Convention all demand that high-risk systems demonstrate safety before deployment. Yet beneath this regulatory consensus lies a critical vacuum: none specifies what ``acceptable risk'' means in quantitative terms, and none provides a technical method for verifying that a deployed system actually meets such a threshold. The regulatory architecture is in place; the verification instrument is not. This gap is not theoretical. As the EU AI Act moves into full enforcement, developers face mandatory conformity assessments without established methodologies for producing quantitative safety evidence - and the systems most in need of oversight are opaque statistical inference engines that resist white-box scrutiny. This paper provides the missing instrument. Drawing on the aviation certification paradigm, we propose a two-stage framework that transforms AI risk regulation into engineering practice. In Stage One, a competent authority formally fixes an acceptable failure probability δ and an operational input domain ε - a normative act with direct civil liability implications. In Stage Two, the RoMA and gRoMA statistical verification tools compute a definitive, auditable upper bound on the system's true failure rate, requiring no access to model internals and scaling to arbitrary architectures. We demonstrate how this certificate satisfies existing regulatory obligations, shifts accountability upstream to developers, and integrates with the legal frameworks that exist today.
Natan Levy, Gadi Perl
School of Computer Science and Engineering The Hebrew University of Jerusalem(HUJI) Jerusalem, Israel · dept. name of organization (of Aff.) The Hebrew University of Jerusalem(HUJI) Jerusalem, Israel