Artificial intelligence systems are increasingly deployed in high impact and safety critical settings, yet security assessment remains difficult to reproduce and defend under audit. Existing approaches often rely on narrative checklists or assessor driven scoring, and they lack an explicit, machine evaluable mapping from observable engineering artefacts to stable technique level outcomes. We present an evidence driven AI security assessment framework that operationalises assessment as a deterministic decision function. The framework normalises heterogeneous artefacts into a project independent Control ID taxonomy scored on a bounded four level ordinal scale, compiles technique level predicates from a pinned MITRE ATLAS snapshot via an explicit mitigation to control mapping, and outputs technique indexed feasibility and impact levels with traceable links back to the triggering evidence. We package all normative choices as a versioned assessment policy object to support repeatable reassessment across snapshots. To ensure semantic correctness, we formally verify boundedness, totality, ordered semantic consistency, and monotonicity of the compiled evaluator over the full declared score domain. We evaluate the framework on five public open source AI projects pinned to explicit repository snapshots, quantify before and after changes under a unified hardening intervention, and validate responsiveness to real engineering changes through fork based implementations of Software Bill of Materials (SBOM) generation and Continuous integration (CI) security scanning gates. Results show consistent downward shifts in feasibility profiles under strengthened observable controls, while worst case residual feasibility persists when technique specific core controls remain absent from the evidence scope.
Figures & tables
Approach
Evidence Grounding
Control Layer
Executable Semantics
Output
Attack Specific
Traceability
System Level
Assessor Dependence
NIST AI RMF 1.0 ( AI, 2023 )
Partial
No
No
Guidance
No
Partial
Broad
High
NCSC ML Security Principles ( National Cyber Security Centre, 2024 )
Partial
No
No
Principles
No
Partial
Broad
High
CSA AI Model Risk Management ( Cloud Security Alliance, 2024 )
Partial
No
No
Gov. artefacts
No
Partial
Governance
High
Microsoft AI Security Risk Assessment
Partial
Partial
Partial
Risk matrix
Partial
Partial
Broad
Med–High
MITRE ATLAS ( MITRE, )
No
No
No
KB
Yes
No
Technique
Medium
Bitton et al ( Bitton et al., 2023 ) .
Partial
Partial
Partial
Scenario score
Partial
Partial
Lifecycle
High
Table 1. Comparison of AI Security Assessment Approaches and the Proposed Framework (KB: knowledge base; MCDA: multi criteria decision analysis; F/I: feasibility/impact).
Figure 1. End to End Workflow of the Evidence Driven Assessment Framework
Level
Anchor definition (consequence severity)
0
Negligible: minimal harm with local and readily reversible effects.
1
Limited: bounded disruption or minor degradation with limited external impact.
2
Major: substantial service degradation, material integrity or confidentiality harm, or costly recovery.
3
Severe: high consequence harm such as prolonged outage, large scale integrity failure, or systemic loss of trust.
Table 2. Ordinal impact rubric (0 lowest, 3 highest) used for technique consequence assignment.
Property
Obligations
Passed
Failed
Pass Rate
Boundedness
86
86
0
100.00%
Totality
54
54
0
100.00%
Ordered Semantic Consistency
108
108
0
100.00%
Monotonicity
54
54
0
100.00%
Overall
302
302
0
100.00%
Table 3. Formal verification summary over the bounded ordinal input domain.
Item
Type
Identifier
Date
KServe
commit
5b033a4
2025-11-3
vLLM
commit
b17039b
2026-01-20
Ultralytics
commit
c529572
2026-01-13
TorchServe
tag
v0.0.9
2024-07-30
cassava-example
tag/commit
-
-
MITRE ATLAS
tag
5.1.1
2025-11-26
Table 4. Pinned Snapshots Used in the Evaluation (Reproducibility Manifest)
Control ID
KServe
vLLM
Ultralytics
TorchServe
cassava-example
api.exposure
2
2
—
1
0
api.access_control
2
1
—
1
0
api.rate_limit
1
0
—
0
0
api.abuse_detection
0
0
—
0
0
sc.version_pinning
3
1
1
0
0
sc.sbom
0
0
0
0
0
Table 5. Baseline Control Scores (0 Weakest, 3 Strongest). ’—’ = not applicable.
Control ID ( c∈C+ )
KServe ( s→s′ )
vLLM ( s→s′ )
Ultralytics ( s→s′ )
TorchServe ( s→s′ )
cassava-example ( s→s′ )
sc.version_pinning
3 → 3
1 → 3
1 → 3
0 → 3
0 → 3
sec.vulnerability_scan
0 → 3
2 → 3
2 → 3
0 → 3
0 → 3
sc.sbom
0 → 3
0 → 3
0 → 3
0 → 3
0 → 3
model.hash_release
1 → 3
1 → 3
1 → 3
1 → 3
0 → 3
org.vdp
2 → 3
2 → 3
2 → 3
1 → 3
0 → 3
org.incident_response
1 → 3
1 → 3
2 → 3
0 → 3
0 → 3
Table 6. Unified Hardening Effects for the Intervention Set C+
Table 7. Raw Data File Sets Used to Bind Each Control ID
Control ID
KServe (A)
KServe (B)
vLLM (A)
vLLM (B)
api.exposure
2
2
2
2
api.access_control
2
1
1
1
api.rate_limit
1
0
0
0
api.abuse_detection
0
1
0
0
sc.version_pinning
3
3
1
1
sc.sbom
0
0
0
0
Table 8. Interrater control scoring comparison on two frozen snapshots. A=Assessor A; B=Assessor B
Project
FISI b
FISI a
Δ FISI
rΔ FISI
WCR b
WCR a
HRS b
HRS a
H b
H a
JSD
KServe
198.00
184.00
14.00
0.0707
6.00
6.00
0.5185
0.5185
3.8366
3.7055
0.0252
vLLM
232.00
210.00
22.00
0.0948
6.00
6.00
0.6852
0.6481
3.8999
3.7372
0.0340
Ultralytics
298.00
256.00
42.00
0.1409
6.00
6.00
0.9815
0.8333
3.9726
3.7978
0.0515
TorchServe
268.00
230.00
38.00
0.1418
6.00
6.00
0.9259
0.8333
3.9531
3.7874
0.0444
Cassava Example
324.00
266.00
58.00
0.1790
6.00
6.00
1.0000
0.8333
3.9890
3.8039
0.0622
Table 9. Aggregate Assessment Metrics Before and After Unified Hardening
Metric
Value
Exact agreement
60.0%
Within one level agreement
96.7%
Cohen’s κ (unweighted)
0.41
Cohen’s κ (linear-weighted)
0.58
Cohen’s κ (quadratic-weighted)
0.73
Table 10. Interrater agreement summary over 30 control instances (15 controls × 2 projects). Weighted κ accounts for ordinal distance on the 0 to 3 scale.
Base Φ
Strict Φ
Permissive Φ
Project
FISIb
FISIa
rΔFISI
FISIb
FISIa
rΔFISI
FISIb
FISIa
rΔFISI
KServe
198.00
184.00
0.0707
324.00
324.00
0.0000
192.00
172.00
0.1042
vLLM
232.00
210.00
0.0948
324.00
324.00
0.0000
230.00
200.00
0.1304
Ultralytics
298.00
256.00
0.1409
324.00
324.00
0.0000
284.00
238.00
0.1620
TorchServe
268.00
230.00
0.1418
324.00
324.00
0.0000
260.00
216.00
0.1692
cassava example
324.00
266.00
0.1790
324.00
324.00
0.0000
324.00
248.00
0.2346
Table 11. Sensitivity of hardening outcomes to alternative policy mapping variants.
Project
Baseline ID
Fork ID
Fork date
KServe
5b033a4
e4ae8f7
2026-02-13
vLLM
b17039b
de2222c
2026-02-13
Table 12. Pinned fork snapshots for concrete control implementation
KServe
vLLM
Control ID
sb
sf
sb
sf
Observable fork evidence (paths/files)
sc.sbom
0
2
0
2
.github/workflows/sbom.yml
sec.vulnerability_scan
0
3
2
3
.github/workflows/codeql.yml
Table 13. Control score changes induced by concrete fork implementations (no manual score edits).
Baseline (upstream)
Fork (implemented)
Project
F=3
F=2
F=1
F=0
F=3
F=2
F=1
F=0
KServe
19
9
24
2
19
9
23
3
vLLM
25
12
17
0
25
12
16
1
Table 14. Technique feasibility distributions before and after concrete fork implementations.
Baseline
Hardened
Project
F=3
F=2
F=1
F=0
F=3
F=2
F=1
F=0
KServe
19
9
24
2
19
9
17
9
vLLM
25
12
17
0
25
10
10
9
Ultralytics
42
11
1
0
38
7
0
9
TorchServe
30
20
4
0
25
20
0
9
Cassava Example
54
0
0
0
43
2
0
9
Table 15. Technique feasibility Distributions Before and After Hardening
Appendix figures & tables1 asset
Supplementary material from the paper’s appendix.
Appendix
Control ID
Meaning
Rubric anchors (0–3)
api.abuse_detection
Identify and suppress query based black box probes
0 : None 1 : Threshold alarm 2 : Multi signal linkage 3 : Automatic handling/isolation
School of Computer Science and Engineering The Hebrew University of Jerusalem(HUJI) Jerusalem, Israel · dept. name of organization (of Aff.) The Hebrew University of Jerusalem(HUJI) Jerusalem, Israel