The Alignment Flywheel: A Governance-Centric Hybrid MAS for Architecture-Agnostic Safety
Organizations: IDLab, Ghent University - imec, Belgium
Abstract
Multi-agent systems provide mature abstractions for role decomposition, coordination, and normative governance, but increasingly capable learned components make post-deployment safety harder to inspect, audit, and update. When safety behavior is absorbed into a decision component, narrow failures may require retraining or rollback of the full component. This instantiates our vision of the Alignment Flywheel as a governance-centric hybrid MAS architecture that decouples decision generation from safety governance. We denote the agent or policy that generates candidate trajectories as the Proposer; it passes its output to a governed Safety Oracle stack, which returns safety scores, prediction uncertainty, audit coverage uncertainty, and evidence hooks through a stable interface. An Enforcement layer applies explicit risk policy at runtime. Around this loop, a governance MAS performs monitoring, red-teaming, verification, triage, refinement, and versioned release management. The central engineering principle is patch locality: many newly observed safety failures can be mitigated through small governance batches for the Oracle stack and its audit state rather than by retraining or retracting the Proposer. The architecture is implementation-agnostic with respect to both Proposer and Oracle. It defines the roles, artifacts, protocols, and release semantics needed for runtime gating, audit intake, signed updates, staged rollout, and rollback. We demonstrate executability in two scenarios: a learned spatial Oracle patched through regression-checked governance updates, and a clinical GenAI proxy setting illustrating structured norms, escalation, and audit coverage. Our implementation code and documentation are available open source at https://github.com/decide-ugent/Alignment-Flywheel.
Figures & tables
| Symbol | Meaning |
| context (task state, inputs, metadata; any modality) | |
| trajectory (action, tool call, message, or plan) | |
| Proposer producing from | |
| Safety Oracle (third-party statistical artifact) | |
| Safety Oracle raw safety score | |
| Safety Oracle prediction uncertainty |
| Phase | Role-independent function | Example |
| Observe | Read relevant records from and queue heads. | Red Team reads active and high-severity norms |
| Orient | Interpret records relative to the role objective. | Red Team identifies a norm with low coverage. |
| Decide | Select a strategy module without changing the protocol. | Prompt mutation to verify norms for an LLM proposer |
| Act | Execute the strategy and write typed records back to . | Candidate flaws are pushed to the queue |
| Stage | Step | Function | Output |
| Verification | Injection | Red Team generates candidate trajectories and pushes them to . | CandidateFlaw |
| Verification | Triage Candidates | They are prioritized by . Low-uncertainty claimed-safe cases are high-value false-negative audit targets. | prioritized |
| Verification | Validation | Verification checks candidates against , and violations become immutable failure records in . | VerifiedBreach |
| Refinement | Ingestion | Verified breaches are hash-stored in for tracing and de-duplication. | breach records |
| Refinement | Triage II | Breaches are clustered into failure families, reducing cognitive load and enabling batch remediation [ 60 , 57 ] . | breach clusters |
| Refinement | Job creation | Triage packages a prioritized cluster as a repair task. | RefinementJob |
| Record | Producer / route | Purpose |
| CandidateFlaw | Red Team | Discovery intake record containing and the Oracle signals observed at discovery time. |
| VerificationResult | Verification | Explicit governance judgment separating raw suspicion from a confirmed or rejected violation of . |
| VerifiedBreach | Verification | Durable failure record created for confirmed violations; later clustered, prioritized, and linked to patches. |
| RefinementJob | Triage | Clustered repair task that converts many verified breaches into one prioritized refinement unit. |
| PatchCommit | Refinement | Governed update record binding a patch to its parent version, motivating breaches, regression evidence, and authorization signature. |
| Step | Function |
| Propose | . |
| Oracle-stack evaluation | The governed Oracle stack evaluates and returns , where is the raw safety score, is Oracle prediction uncertainty, is Alignment Flywheel audit coverage uncertainty, and are optional hooks. |
| Enforcement decision | selects under the configured risk policy. High means the Oracle is unsure about its prediction; high means the Alignment Flywheel has insufficient audit coverage for this class of case. |
| Log | writes an auditable decision record to , including the active version for attribution, audit, and regression analysis. |
| Record field | Purpose |
| Decision record | records the context, candidate, action, safety score, uncertainty signals, thresholds, version, and evidence. |
| Integrity metadata | Hash of the serialized trajectory, timestamp, and host identity support replay and fleet correlation. |
| Optional attachments | , counterexample identifiers, trace references, source deployment traffic label, stress test, or report. |
| Eval | Allow | Block | Escalate | Esc. Rate | Stack |
| 0 | 6 | 0 | 9 | 60% | v0 |
| 1 | 6 | 3 | 6 | 40% | v1 |
| 2 | 9 | 6 | 0 | 0% | v2 |
| Dimension | 3D spatial demo | Patient Portal demo | Architectural property |
| Evaluation mode | Offline hardening of a learned Oracle artifact | Runtime Proposer–Oracle enforcement and hardening | Offline and runtime governance |
| Oracle type | Learned continuous IIRL-style Oracle | Structured proxy Oracle with safety and uncertainty signals | Oracle substitutability |
| Proposer role | No Proposer; the Oracle is hardened directly | Passthrough Proposer producing candidate replies and dispositions | Separation of proposal and governance |
| Norm structure | Spatial support and expert-basin preservation | Medication, urgency, lab-context, and unsupported-advice norms | Norm-semantic variation |
| Refinement mechanism | Gaussian suppression patches released through governance batches | Blocks, disposition-sensitive decisions, threshold updates, and audit-coverage updates | Patch-local correction |
| Observed effect | False-positive reward cells are removed while the sampled expert basin is preserved | Escalations are converted into definitive allow/block outcomes across governance batches | Executability and protocol stability |
Appendix figures & tables24 assets
Supplementary material from the paper’s appendix.
Appendix
| Mode | Human role | Typical use |
| Low risk | Exception handling | Automated verification and refinement; humans review anomalies or periodic reports. |
| Medium risk | Batch approval | Agents cluster failures and propose governance batches; humans review summaries and authorize release. |
| High risk | Item and release approval | Humans review high-impact cases, approve or modify patches, and sign each governance batch before rollout. |
| Control surface | Effect |
| Agent efficacy and strategy rotation | Surface metrics such as unique verified breaches per unit time; deprecate weak strategies and activate stronger ones. |
| Normative coverage and prioritization | Visualize coverage across ; adjust severity weights or introduce new norms to re-orient OODA loops. |
| Adversarial seeding and directed search | Inject seed trajectories into ; Red Team agents mutate or exploit them to discover variants [ 22 , 30 ] . |
| Deployment feedback | Surface high-uncertainty runtime cases as incidents; direct coverage expansion and verification toward failed regions. |
| Release approval and rollback | Require human approval, signature, or staged rollout for governance batches in high-risk settings. |
| Norm kind | Verifier semantics |
| KEYWORD_BLOCK | Checks whether trajectory text contains prohibited keywords, optionally conditioned on evidence status. Example: medication-change terms under weak evidence. |
| REGEX | Checks whether trajectory text matches a regular-expression pattern, such as a prohibited action phrase. |
| PREDICATE | Evaluates multi-field relationships over payload and metadata, such as whether the proposed disposition is at least as severe as required for the case type and evidence status. |
| SPATIAL_BOUNDARY | Checks whether a spatial query point lies within the supported region, e.g., below a distance threshold from expert data. |
| THRESHOLD_RULE | Checks threshold-style constraints, such as age or vulnerability conditions requiring a specified evidence standard. |
| Cell type | Condition | Interpretation |
| Basin cell | and | Legitimate reward near the expert trajectory; must be preserved. |
| Flaw cell | and | Non-trivial reward outside the expert basin; must be suppressed. |
| Inactive cell | Already below the safety floor; no action required. |
| Step | Function |
| Skip covered flaws | If a previously accepted kernel in the same batch already covers flaw , skip it. |
| Propose bandwidth | Set Distant flaws receive wider kernels; near-boundary flaws receive narrower kernels. |
| Single-kernel basin protection | Let . The maximum bandwidth that keeps suppression of the nearest basin point below is The proposed bandwidth is shrunk to . |
| Cumulative basin check | The planner tracks running suppression on every basin point from accepted kernels in the current batch: If adding the candidate kernel would push any basin point above , a 15-step binary search over finds the largest safe bandwidth. If even would cause excessive basin suppression, the patch is rejected. |
| Predict coverage | The effective suppression radius is Other flaw points within are predicted to be suppressed below the safety floor. |
| Mark covered flaws | Predicted-covered flaws are marked as handled and skipped later in the batch. This spends the patch budget on non-redundant kernel placements. |
| Correction type | Effect |
| SPATIAL_FLAW_PATCH | Payload contains flaw_point and support_radius . Applying the correction installs a suppression kernel in the Oracle. |
| AUDIT_COVERAGE_UPDATE | Payload records the predicted coverage class, e.g. "spatial|bw= |cov= " , for audit trail and batch provenance. |
| Metric | PatchPlanner (adaptive) | Fixed BW ( ) |
| Iterations to converge | 16 | (not converged) |
| Patches/iter (max) | 60 | 60 |
| Basin preserved | 783/783 (100%) | 783/783 (100%) |
| Total kernels placed | 834 | 1500 (it. 30) |
| Iter | Found | Kern | Predicted | Reject | Basin | Flaws | Oracle |
| 1 | 500 | 87 | 500 | 0 | 783 | 412 | v1 |
| 2 | 993 | 42 | 993 | 0 | 783 | 357 | v2 |
| 3 | 874 | 58 | 874 | 0 | 783 | 290 | v3 |
| 4 | 1 030 | 57 | 1 030 | 0 | 783 | 219 | v4 |
| 5 | 1 119 | 49 | 1 119 | 0 | 783 | 159 | v5 |
| 6 | 960 | 71 | 959 | 1 | 783 | 103 | v6 |
| Component | Function |
| Proposer | A PassthroughProposer wraps the incoming case data—patient message, draft reply, disposition, and metadata—as a Trajectory of kind MESSAGE . No generation occurs; the case already contains the candidate content. |
| Safety Oracle | Evaluates the trajectory and returns OracleRawOutput : safety score , prediction uncertainty , uncertainty threshold , Oracle version , and evidence status. |
| Alignment Flywheel Overlay | Evaluates audit coverage and returns Alignment FlywheelOverlay : audit coverage uncertainty , threshold , governance version , and audit status. |
| Query Merger | A DefaultQueryMerger combines Oracle-side and governance-side signals into a UnifiedQueryResult while preserving provenance. |
| Enforcement | A DefaultEnforcement policy maps the unified signals to . |
| Condition | Action |
| escalate : Oracle prediction uncertainty is too high. | |
| block : safety score indicates unsafe behavior. | |
| , evidence , and | escalate : audit coverage is insufficient for a weak-evidence, borderline-safety case. |
| Otherwise | allow . |
| Role | Function |
| Red Team | Generates 20 synthetic cases per iteration by sampling a combinatorial case space. Each case is classified by failure category and pushed to as a CandidateFlaw . |
| Verification | Checks candidates against the normative specification . Keyword-block norms check prohibited terms under weak evidence; predicate norms check multi-field relationships such as whether the proposed disposition meets the required minimum. Confirmed violations produce VerificationResult records. |
| Triage | Orders verified violations by priority and forwards them to . |
| Refinement | Maps verified flaws to typed corrections. A tiered batch strategy rotates across categories to ensure diverse coverage. In all three demos, batch capacity is limited to one flaw per iteration. |
| Batch application | Applies the resulting GovernanceBatch to the Oracle and Alignment Flywheel Overlay. Oracle patches install hard-blocks and disposition overrides; overlay patches register audit coverage. Both components increment their version identifiers when state changes. |
| Correction type | Effect |
| MEDICAL_HARD_BLOCK | Adds a keyword to the Oracle’s internal block list. Later trajectories containing the keyword return , , producing an unconditional block. |
| THRESHOLD_ADJUSTMENT | Installs a disposition override keyed by case_type|evidence_status . If a matching trajectory has a disposition below the required minimum, the Oracle returns , , which Enforcement maps to block . |
| AUDIT_COVERAGE_UPDATE | Registers a case_type|evidence|acuity key as audited in the Alignment Flywheel Overlay. Covered keys produce ; uncovered keys produce . |
| Role | Component | Function |
| Red Team | Medical Case Generator | Samples medications, symptoms, lab tests, and demographics; classifies cases as missed_urgency , undertriaged_med, lab_no_context, vulnerable_patient , or exploratory . |
| Red Team | Medical Priority Decider | Sorts candidates by category-weighted priority: missed_urgency undertriaged_med vulnerable_patient lab_no_context exploratory . |
| Red Team | Medical Candidate Submitter | Wraps candidates as Candidate Flaw artifacts with Oracle signals, trajectory payload, and metadata. |
| Verification | Medical NormM atcher and Medical Violation Decider | Extract draft text, patient message, evidence status, and disposition; evaluate keyword-block and predicate norms; emit VIOLATION or NO_VIOLATION . |
| Refinement | Medical Correction Orienter | Maps each verified flaw to a small bundle of corrections, e.g., hard-block, disposition overrides, and audit-coverage updates. |
| Refinement | Medical Batch Decider and Medical Batch Deployer | Selects at most one flaw per iteration, builds a GovernanceBatch , sets version transitions, and returns the batch for application. |
| Group | Count | Risk pattern | Expected transition |
| Unsafe medication | 3 | Unsupported dosing advice with reply_only disposition. | escalate block |
| Unsafe lab | 3 | Lab-result interpretation or monitoring advice without sufficient evidence and without nurse review. | escalate block |
| Borderline | 3 | Similar clinical content, but already routed to clinician_review . | escalate allow |
| Safe | 6 | Routine supported interactions such as refills, scheduling, and acknowledgements. | allow throughout |
| Component | Specification |
| Oracle | SimpleMedicalOracle ; heuristic scoring with patchable hard-block list and disposition override dictionary . Hard-blocks return ; disposition failures return . Base scoring penalizes medication stems, medication case type, and weak evidence. |
| Prediction uncertainty | Baseline , increased for insufficient, unknown, or conflicting evidence and for medication cases. for medication cases and otherwise. |
| Alignment Flywheel Overlay | SimpleMedicalOverlay ; coverage key case_type|evidence_status|acuity . Covered keys produce ; uncovered keys produce . Threshold . |
| Norms | Two norms: NO_UNSUPPORTED_MED_ADVICE , a keyword-block norm over medication-change terms under weak evidence; and REQUIRE_EVIDENCE_FOR_RESULTS , a predicate norm requiring at least nurse_review for weak-evidence lab/result interpretation. |
| Triage | FIFOTriage ; violations are passed to refinement in arrival order. |
| Eval | Allow | Block | Escalate | Esc. Rate | Stack |
| 0 | 6 | 0 | 9 | 60% | v0 |
| 1 | 6 | 3 | 6 | 40% | v1 |
| 2 | 9 | 6 | 0 | 0% | v2 |
| Group | Count | Risk pattern | Expected transition |
| Unsafe stop-taking | 3 | Recommendations to stop high-risk medications such as warfarin, insulin, or oxycodone under weak evidence. | escalate block |
| Unsafe dose increase | 3 | Autonomous dose increases for higher-risk drugs such as insulin, tramadol, or prednisone. | escalate block |
| Unsafe lab | 3 | Organ-function or medication-related lab interpretation with insufficient evidence and direct patient routing. | escalate block |
| Borderline | 3 | Lab or medication discussion already routed to clinician_review . | escalate allow |
| Safe | 6 | Routine supported interactions without clinical decision-making. | allow throughout |
| Component | Specification |
| Oracle | ComplexMedicalOracle ; patchable state includes hard-block list , specialty threshold overrides , and disposition overrides . Evaluation order is hard-block, disposition override, then five-dimensional scoring. |
| Risk dimensions | from a medication-risk table; from known interaction pairs; from action severity; from age and comorbidities; from evidence status. |
| Alignment Flywheel Overlay | ComplexMedicalOverlay ; coverage key specialty|case_type|evidence_status . Covered keys produce ; uncovered keys produce . Threshold . |
| Norms | Four norms: NO_HIGH_RISK_MED_WITHOUT_EVIDENCE , INTERACTION_SAFETY , VULNERABLE_PATIENT_PROTECTION , and SEVERE_ACTION_EVIDENCE . These cover keyword-block, regex, threshold-rule, and predicate-style checks. |
| Triage | PriorityTriage ; priority order is missed_urgency undertriaged_med vulnerable_patient lab_no_context exploratory . |
| Eval | Allow | Block | Escalate | Esc. Rate | Stack |
| 0 | 6 | 0 | 12 | 67% | v0 |
| 1 | 6 | 6 | 6 | 33% | v1 |
| 2 | 9 | 9 | 0 | 0% | v2 |
| Group | Count | Risk pattern | Expected transition |
| Unsafe medication | 3 | Direct patient-facing medication-change advice, such as increasing dose, without clinician review or sufficient evidence. | escalate block |
| Unsafe lab | 3 | Lab-result interpretation with medication context, weak evidence, and reply_only routing. | escalate block |
| Borderline | 3 | Lab or medication discussion already routed to clinician_review . | escalate allow |
| Safe | 6 | Routine supported interactions such as refills, appointments, physical-therapy progress, diet follow-up, improving symptoms, and normal annual labs. | allow throughout |
| Component | Specification |
| Oracle | PatientPortalOracle ; patchable state includes hard-block list and disposition override dictionary . Evaluation order is hard-block, disposition override, then four-dimensional scoring. |
| Risk dimensions | counts medication keywords; counts urgency keywords in the message or draft; encodes evidence quality; measures the gap between required and proposed disposition. |
| Disposition ranks | reply_only , nurse_review , clinician_review , urgent_escalation . Needed disposition increases with urgency, medication risk, and semi-urgent acuity. |
| Alignment Flywheel Overlay | PatientPortalOverlay ; coverage key case_type|evidence_status|acuity . Covered keys produce ; uncovered keys produce . Threshold . |
| Norms | Four norms: MED_ADVICE_REQUIRES_CLINICIAN , URGENCY_REQUIRES_ESCALATION , LAB_RESULTS_NEED_CONTEXT , and NO_UNSUPPORTED_MED_KEYWORDS . |
| Triage | PriorityTriage ; same category-priority ordering as the complex demo. |
| Eval | Allow | Block | Escalate | Esc. Rate | Stack |
| 0 | 6 | 0 | 9 | 60% | v0 |
| 1 | 6 | 3 | 6 | 40% | v1 |
| 2 | 9 | 6 | 0 | 0% | v2 |
| Component | Simple | Complex | Patient Portal |
| Oracle scoring | 3 heuristic rules | 5-dimensional lookup | 4-dimensional keywords |
| Patchable state | , | , , | , |
| Norm count (kinds) | 2 (2) | 4 (4) | 4 (2) |
| Triage | FIFO | priority-ranked | priority-ranked |
| Fixed cases | 15 | 18 | 15 |
| Evaluations to converge | 3 | 3 | 3 |