Reliability Engineering for AI Systems: Challenges, Methods, and Directions
Organizations: Arizona State University, Tempe, AZ, USA · Virginia Tech, Blacksburg, VA, USA · City University of Hong Kong, Hong Kong
Abstract
AI reliability concerns whether an AI system performs its intended function dependably over a stated period and under stated operating conditions, with stated evidence. As these systems become more autonomous, that function includes more than a correct output. Retrieval, memory, tool use, permissions, human oversight, and interactions among systems must operate consistently and safely, and, for generative systems, so must the reasoning process that produces the output. Average benchmark accuracy measures capability; it does not quantify this broader reliability claim. This paper adapts established reliability engineering methods, from failure definitions and operational envelopes to FMEA, accelerated testing, field monitoring, and reliability growth, to AI systems. A four-level diagnostic framework classifies failures as component, operational-loop, agentic-conduct, or network and governance failures. Test, evaluation, verification, and validation (TEVV), sequential monitoring, and FRACAS create and refresh evidence. SMART provides statistical guidance for measurement, analysis, assessment, and test planning; the NIST AI Risk Management Framework provides organizational guidance for governance, evaluation, monitoring, and mitigation. Three cases illustrate the program: adversarial testing of a convolutional neural network, perception-error propagation, and autonomous-vehicle disengagements. Established reliability engineering provides a usable foundation; new measurements and safety guardrails are still needed as these systems are self-evolving.
Figures & tables
| Challenge | Reliability-engineering response (section) | Research / practice direction |
|---|---|---|
| No agreed failure definition | Failure, envelope, and evidence stated jointly; level-aware FMEA (§ 2 , § 3.2 ) | Standard AI failure taxonomies with severity classes usable in FRACAS |
| Averages measure capability, not a time-indexed claim | Consistency, robustness, predictability, safety, recoverability, governability; all-of- reporting (§ 4 ) | Benchmarks that report process soundness, trajectories, tails, and severity, not means |
| Exposure is rarely recorded, so counts are not rates | Declare the data type, then model it; mileage offset in the AV case (§ 4 , § 6 ) | Registries that publish deployed population, task-hours, or tool calls |
| Overdispersion and dependence in field records | Quasi-Poisson inflation; recurrent-event and triggering intensities (§ 6 ) | Dispersion and clustering models for agent and fleet event streams |
| Nonstationarity: models, corpora, and tools change | Sequential change detection; revalidation triggers; change control (§ 4 ) | Reliability growth with model-version and corpus covariates |
| Agents act, remember, and can breach containment | Information quality, memory governance, action discipline; action gating (§ 5 ) | Event logs with tool, memory, and policy-gate covariates |
| Classical method | Application to AI | What must be reinterpreted |
|---|---|---|
| Time-indexed reliability | Intended function under a stated envelope for a stated period | The function includes reasoning traces, tools, retrieval, and actions, not only a label |
| FMEA / FRACAS | Level-aware AI-FMEA; incident taxonomy with corrective action | Modes at interfaces, memory, and conduct; narratives need denominators |
| Accelerated testing | Use-rate, input-data (adversarial, error injection), and environment acceleration | Stress is perturbation, imbalance, and trap content, not temperature |
| SPC / change detection | Score calibrate sequential test on residuals or embeddings | Streaming dependence breaks naive exchangeability |
| Recurrent events / NHPP | Disengagements, misclassifications, module errors with an offset for exposure | Exposure is miles, task-hours, or tool calls; field series are overdispersed |
| Reliability growth / SRGM | Retraining, policy revision, and gate tightening as the “fix” process | A software or policy change is the repair action |
| Level and failure type | Representative failure mechanisms | Representative controls |
|---|---|---|
| 1. Technical component failure | Incorrect prediction, unsupported generation, unsound or unfaithful reasoning, miscalibration, covariate or concept drift, latency, schema breakage, and retrieval, software, infrastructure, or API/tool failure | Validation, calibration, drift monitoring, source verification, process soundness checks |
| 2. Operational loop failure | Wrong tool selection, stale or injected memory, skipped validation, runaway retries, late escalation, and failed stopping or sandbox containment checks | State-machine orchestration, validators, governed memory, Transactional No-Regression, rollback, escalation |
| 3. Agentic conduct failure | Violations of authorization, honesty, privacy, safety, reversibility, or oversight; unauthorized action and scope violation | Role-specific SOPs, least privilege, independent policy gates, kill switches |
| 4. Network and governance failure | Cascade- or hub-amplifying structures, compressed handoffs, misaligned incentives, dependence, weak aggregation, long-path failure, and correlation failure | Topology design, independent verification, dissent preservation, circuit breakers |
| Failure mode | Lvl | Typical causes | Detection signals | Engineering controls |
|---|---|---|---|---|
| Distribution / concept drift | 1 | New users, sensors, policies, corpora | Embedding drift; sentinel-case error; decay with stable inputs | Drift monitoring; targeted reevaluation; conservative routing |
| Brittleness to perturbation | 1–2 | Surface-cue reliance; fragile prompts; schema sensitivity | Large delta under paraphrase or tool fault | Structured I/O; schema validation; safe-mode fallbacks |
| Miscalibration / overconfidence | 1 | Objective mismatch; weak uncertainty | High-confidence errors; calibration gaps | Calibration; selective automation; abstain/escalate |
| Hallucination / unsupported generation | 1 | Weak grounding; poor provenance | Missing citations; source contradictions | Evidence-first workflows; provenance checks; validators |
| Unsound reasoning | 1–2 | Unjustified leaps; invented premises; post-hoc traces; omitted constraints | Step-verifier disagreement; plan–execution mismatch; contradiction across repeats | Process checks before action; grounding gates; abstain or escalate when the reasoning chain is not safe |
| Unsafe actions | 2–3 | Ambiguous SOP; excess permissions; mistaken target; no rollback | Scope mismatch; credential access; destructive-call attempt | Least privilege; independent action gate; confirmation; rollback |
| Dimension | Representative measurement questions |
|---|---|
| Consistency | If we run the same task multiple times, do we get the same outcome and similar trajectory? How often do all repeated trials succeed (all-of- ), including correct persistent state and no collateral effects? ( Raj et al., 2026 ; Li et al., 2026 ) Do key intermediate claims in a reasoning trace agree across repeats? |
| Robustness | How does performance change under paraphrases, formatting changes, tool failures, schema changes, or mild shifts? Does the system degrade smoothly or collapse? |
| Predictability | When the system is confident, is it correct (calibration)? Can confidence separate success from failure (discrimination)? Can it abstain appropriately? Does stated confidence track process soundness, not only the final answer? |
| Safety | What is the violation probability? When violations occur, how severe are they, and what fraction are blocked before an external effect? |
| Recoverability | Can failures be detected, stopped, reversed, or contained? What are MTTD/MTTR and rollback/no-regression success? ( Chen et al., 2025 ) |
| Governability | Can decisions be audited and constrained over time? Are lineage and deletion obligations executable? ( Kumar et al., 2026b ) |
| Element | This case | Evidence |
|---|---|---|
| Intended function | Public-road autonomous driving under tester rules | CA AVT program |
| Operational envelope | CA public roads, Dec. 2017–Nov. 2019; Waymo, Cruise, Pony AI, Zoox | DR-AIR event and mileage files |
| Failure definition | Disengagement (exit from autonomous mode); not a crash | Event CSV; SOTIF caveat ( International Organization for Standardization, 2022 ) |
| Metrics | Events per mile; and (growth); Pearson | Table 7 ; Figure 3 |
| TEVV / data | Mandated field reporting (not a laboratory demo) | Manufacturer-month GLM (this paper); vehicle-level NHPP in Min et al. (2022) |
| Monitoring | Monthly intensity vs mileage exposure | Manufacturer-month aggregation |
| Mfr. | Events | Per mi | ||||
|---|---|---|---|---|---|---|
| Waymo | 224 | 0.083 | 2.03 | 0.016 | 0.979 | |
| Cruise | 154 | 0.120 | 1.09 | 0.012 | 0.930 | |
| Pony AI | 43 | 0.225 | 2.99 | 0.042 | 0.864 | |
| Zoox | 58 | 0.593 | 1.12 | 0.023 | 1.008 |
| DR-AIR case | Data type | Level | Model / decision |
|---|---|---|---|
| Case A: AV disengagements | Recurrent events; mileage offset | 1–2 | Quasi-Poisson GLM (Table 7 ); vehicle-level NHPP ( Min et al., 2022 ) |
| Case B: CNN + FGSM/PGD ( Faddi et al., 2025 ) | Failure counts; accuracy trajectory | 1 | Covariate SRGM + resilience |
| Case C: perception modules ( Pan et al., 2024 ) | Module recurrent errors | 1, 4 | Triggering NHPP; virtual ALT with known ground truth |
| AI incident narratives | Text; no exposure | 1, 3 | Taxonomy only; cannot estimate rates |
| LLM agents / MAS | Not in DR-AIR | 2–4 | Need event logs with exposure and tool, memory, and topology covariates |