The EU AI Act introduces extensive compliance requirements for organizations that develop, deploy, or integrate AI systems. Many of these requirements are directly relevant to security and privacy, while also addressing closely related issues such as data governance, transparency, accuracy, and robustness. However, stakeholders such as small-to-medium businesses and individual developers often lack the legal expertise required to interpret these obligations and translate them into engineering and governance practices. This disconnect creates challenges for implementing the EU AI Act and may lead to missing safeguards or misdirected development and deployment efforts. To address this, various automated EU AI Act compliance checkers (AIACCs) have emerged, claiming to streamline compliance assessments and provide practical guidance. In this paper, we present the first empirical study and assessment of AIACCs. We characterize 12 mainstream AIACCs across multiple dimensions, evaluate their legal coverage and alignment, and analyze checker-generated compliance reports for structure, determinacy, and actionability. We find that the quality of AIACCs varies significantly and that they currently can only serve as early-stage orientation tools. Specifically, we observe inconsistent interaction modes and user-friendliness, a tendency to overly simplify or omit key obligations, and a failure to provide determinate, actionable guidance. As a result, reliance on the current generation of AIACCs may foster a false sense of compliance. With our study, we provide a critical baseline of the current AIACC landscape. We further offer design principles for the implementation of more reliable compliance-support tools.
Figures & tables
Fig. 1 : Examples (C1, C8, C9, and C12) of compliance checkers. (a.1) The multiple-choice question in the dynamic-branching questionnaire mode. (a.2) The branching logic of the questionnaire. (b.1) The polar question in the fixed questionnaire mode. (c.1) Conversation regarding cybersecurity obligations in the conversational assistant mode. (d.1) The interactive visualization of AI Act concepts in technical analysis tool mode. (d.2) The assessment of AI model documentation.
Branch
Scenario
Role
S1
Prohibited
Workplace emotion recognition
Deployer
S2
High-risk
Recruitment CV screening
Provider
S3
Transparency
E-commerce chatbot
Deployer
S4
GPAI
GPAI API provider
Provider
S5
Minimal risk
Email spam filter
Deployer
S6
Out of scope
Research-only AI model
Provider
TABLE I : Representative scenarios used for AIACC-generated report collection.
Fig. 2 : Selection process of compliance checkers.
#
Name
Provider Sector
Registration
Mode
AI-empowered
1
EU AI Act Compliance Checker [ 41 ]
Non-profit organization
Not required
Dynamic‑Branching Questionnaire
✗
2
TRAIL ML EU AI Act Compliance Checker [ 44 ]
Industry provider
Not required
Dynamic‑Branching Questionnaire
✗
3
TÜV Risk Navigator [ 45 ]
Industry provider
Not required
Dynamic‑Branching Questionnaire
✗
4
Statworx AI Act Quick Check [ 46 ]
Industry provider
Required
Dynamic‑Branching Questionnaire
✗
5
AI Act Conformity Check [ 47 ]
Non-profit organization
Required
Dynamic‑Branching Questionnaire
✗
6
EU AI Act Compliance Checker [ 43 ]
Public authority
Not required
Dynamic‑Branching Questionnaire
✗
TABLE II : Compliance checkers selected for our study.
Mode
DBQ
FQ
CA
TAT
# of AIACC
1
2
3
4
5
6
7
8
9
10
11
12
Risk Classification
●
●
●
●
●
●
○
●
●
●
○
○
Article Mapping
●
◐
●
●
●
●
◐
○
◐
◐
●
◐
User Guidance
●
●
●
◐
◐
●
○
◐
●
●
○
○
Examples
○
○
○
○
○
◐
●
●
◐
○
○
○
Interactive Q&A
○
○
○
○
○
○
○
○
●
○
○
○
TABLE III : The characterization of 12 AIACCs. ●: high-level support; ◐: intermediate support; ○: low-level support. “DBQ” refers to dynamic branching questionnaire; “FQ” refers to fixed questionnaire; “CA” refers to conversational assistant; “TAT” refers to technical analysis tool.
Fig. 3 : Different results using different base models under identical prompt. GPT-4o produced incorrect result (left box) while GPT-5 produced correct result (right box).
Fig. 4 : A compliance checking conversation flow indicating inconsistency and user-pleasing behavior. The checker first produced a misclassified prohibited answer, then changed to high-risk after user objection.
Fig. 5 : Examples of readability issues in AIACCs.
Mode
DBQ
FQ
CA
TAT
# of AIACC
1
2
3
4
5
6
7
8
9
10
11
12
Input
System Description
○
○
○
○
●
○
○
○
◐
○
●
●
User’s Contact Information
○
◐
○
●
●
○
○
●
○
◐
○
○
AI Model or System
●
●
●
●
○
●
○
●
◐
●
○
○
Domain of Application
●
●
●
●
●
●
●
●
●
●
○
○
TABLE IV : The breakdown of required inputs and generated outputs of AIACCs.
Fig. 6 : Benchmark score generated by C11 for the Cyberattack Resilience requirement in AI Act Articles 15 & 55. Scores are calculated based on the benchmarking results.
Metric
C1
C2
C3
C4
C5
C6
C7
C8
Avg. #Words
181.0
68.2
119.7
46.3
303.0
335.2
1023.3
324.8
Avg. #Sent.
11.7
4.2
12.7
3.3
21.0
16.7
47.7
31.5
TABLE V : Basic statistics of AIACC-generated reports.
Section Category
C1
C2
C3
C4
C5
C6
C7
C8
Role determination
✗
✓
✓
✗
✗
✓
✗
✓
Risk-level assessment
✓
✓
✓
✓
✓
✓
✓
✓
Object determination
✗
✗
✓
✗
✗
✓
✗
✗
Legal justification
✗
✓
✓
✓
✓
✗
✓
✗
Obligations
✓
✓
✓
✓
✓
✓
✓
✓
AI literacy obligations
✓
✓
✗
✗
✗
✓
✗
✗
TABLE VI : Composition of checker-generated reports. *Disclaimers may appear outside of the report (e.g., within text on the webpage). We mark “✓” whenever disclaimers are found, regardless of the location.
Fig. 7 : Category composition of checker-generated reports by checker and scenario.
Fig. 8 : Average Hedge Sentence Rate by checker across the six scenarios. Higher values indicate greater hedge-cue prevalence and weaker linguistic determinacy.
Fig. 9 : Average Keyword Loss of Specificity by checker across generated reports. Lower values indicate that directive keywords more often correspond to direct instructions.
Mode
DBQ
FQ
CA
TAT
(#)
1
2
3
4
5
6
7
8
9
10
11
12
General
Art. 4 (AI Literacy)
✓
✓
✗
✗
✗
✓
✗
✗
✓
✓
✗
✓
Art. 5 (Prohibited AI Practices)
✓
✓
✓
✓
✓
✓
✓
✓
✓
✓
✗
✗
High-Risk AI
Art. 6 (High-Risk Classification)
✓
✓
✓
✓
✓
✓
✓
✓
✓
✓
✗
✗
TABLE VII : Coverage of EU AI Act Articles across AIACCs.
Fig. 10 : Different wordings of “real-time biometric identification” question across checkers C1-C3.
Scenario
C1
C2
C3
C4
C5
C6
C7
C8
Avg.
S1
0.333
0.667
0.333
1.000
1.000
1.000
0.690
0.342
0.671
S2
0.500
0.800
0.200
0.750
0.300
0.417
0.655
0.333
0.494
S3
0.833
0.750
0.429
0.750
0.571
0.941
0.737
0.409
0.678
S4
0.522
0.800
0.143
1.000
0.600
0.611
0.680
0.395
0.594
S5
1.000
0.750
0.429
1.000
0.571
1.000
0.680
0.409
0.730
S6
0.333
0.750
0.500
1.000
1.000
1.000
0.719
0.571
0.734
TABLE VIII : Hedge Sentence Rate ( ↓ ) by scenario and checker.
Scenario
C1
C2
C3
C4
C5
C6
C7
C8
Avg.
S1
NA
NA
NA
NA
0.000
1.000
0.563
0.714
0.569
S2
0.833
1.000
NA
1.000
0.250
1.000
0.588
0.750
0.775
S3
1.000
1.000
1.000
1.000
1.000
1.000
0.700
0.750
0.931
S4
1.000
1.000
NA
1.000
1.000
1.000
0.600
0.800
0.914
S5
1.000
1.000
1.000
1.000
0.000
1.000
0.600
0.750
0.794
S6
NA
NA
1.000
1.000
1.000
1.000
0.700
1.000
0.950
TABLE IX : Keyword Loss of Specificity ( ↓ ) by scenario and checker. We report “NA” for the report that has no keyword-containing sentences.
The EU AI Act (Regulation 2024/1689) imposes technical obligations on high-risk AI providers, yet Articles 8-15 were drafted for predictive AI and leave seven technical gaps when applied to generative systems, spanning non-deterministic data governance, training-data provenance, continuous conformity, human oversight, open-ended robustness, emergent risk, and generative fairness. We deliver Governance-as-Code (GaC), a framework of 43 machine-checkable acceptance criteria across six compliance modules that run in a CI/CD pipeline and emit Article-indexed audit evidence, and we show the actual Rego policy code rather than merely describing it. Our central commitment is that the Act's open-textured standards ("appropriate levels," "possible biases") become declared, auditable numbers: robustness thresholds are derived from the provider's documented baseline and a state-of-the-art floor, and framing bias is collapsed into eight measurable proxies tested by counterfactual demographic probing. We also correct who owes what, since under Article 25 and Chapter V a downstream deployer relies on the upstream provider's Article 53 training-data summary and documents only the layers it controls, so GaC verifies that summary rather than demanding per-sample documentation the deployer never had. We validate on two enterprise deployments, a high-risk advisory chatbot and a limited-risk content generator, benchmarking against a manual expert audit rather than documentation artifacts that were never designed to enforce compliance. GaC reproduces all of the manual audit's findings, including three penalty-triggering violations, while cutting audit labor by roughly 75%.
Large language models now produce legal text of at least median quality, yet no existing benchmark can evaluate whether they perform doctrinal legal reasoning, which forms the interpretive core of legal work, rather than the ancillary, paralegal tasks that most current legal-AI evaluations measure. This measurement gap is not only methodological but legal: the EU AI Act makes "appropriate accuracy" a binding requirement for high-risk AI used in the judicial domain, yet that requirement cannot acquire operational content without the very doctrinal-reasoning benchmark the field lacks.
Michèle Finck
Chair of Law and Artificial Intelligence and Director, CZS Institute for Artificial Intelligence and Law, University of Tübingen.
The systematic assessment of AI systems is increasingly vital as these technologies enter high-stakes domains. To address this, the EU's Artificial Intelligence Act introduces AI Regulatory Sandboxes (AIRS): supervised environments where AI systems can be tested under the oversight of Competent Authorities (CAs), balancing innovation with compliance, particularly for startups and SMEs. Yet significant challenges remain: assessment methods are fragmented, tests lack standardisation, and feedback loops between developers and regulators are weak. This paper operationalises the AIRS lifecycle. We map the sandbox journey into 29 concrete activities, from pre-participation guidance through application, preparation, participation, exit, and post-participation monitoring, and we distinguish between a Core AIRS centred on regulatory oversight and an Extended AIRS that additionally embeds structured technical testing through an AI Technical Sandbox (AITS). From this mapping we derive 15 infrastructural and governance requirements that an AITS must satisfy, each linked to the activities it supports and, for high-risk systems, to the provider obligations set out in Articles 9-15 of the AI Act. The framework aims to address multiple stakeholders: CAs gain structured workflows for applying legal obligations; technical experts can integrate robust evaluation methods; and AI providers access a transparent pathway to compliance. We conclude by outlining the Sandbox Configurator, an open-source framework intended to instantiate AITS environments from these requirements, and by discussing how a shared technical foundation can support a scalable and innovation-friendly European infrastructure for trustworthy AI governance.