cs.MAFeb 28, 2026

The Alignment Flywheel: A Governance-Centric Hybrid MAS for Architecture-Agnostic Safety

Authors: Elias Malomgré, Pieter Simoens

Organizations: IDLab, Ghent University - imec, Belgium

Abstract

Multi-agent systems provide mature abstractions for role decomposition, coordination, and normative governance, but increasingly capable learned components make post-deployment safety harder to inspect, audit, and update. When safety behavior is absorbed into a decision component, narrow failures may require retraining or rollback of the full component. This instantiates our vision of the Alignment Flywheel as a governance-centric hybrid MAS architecture that decouples decision generation from safety governance. We denote the agent or policy that generates candidate trajectories as the Proposer; it passes its output to a governed Safety Oracle stack, which returns safety scores, prediction uncertainty, audit coverage uncertainty, and evidence hooks through a stable interface. An Enforcement layer applies explicit risk policy at runtime. Around this loop, a governance MAS performs monitoring, red-teaming, verification, triage, refinement, and versioned release management. The central engineering principle is patch locality: many newly observed safety failures can be mitigated through small governance batches for the Oracle stack and its audit state rather than by retraining or retracting the Proposer. The architecture is implementation-agnostic with respect to both Proposer and Oracle. It defines the roles, artifacts, protocols, and release semantics needed for runtime gating, audit intake, signed updates, staged rollout, and rollback. We demonstrate executability in two scenarios: a learned spatial Oracle patched through regression-checked governance updates, and a clinical GenAI proxy setting illustrating structured norms, escalation, and audit coverage. Our implementation code and documentation are available open source at https://github.com/decide-ugent/Alignment-Flywheel.

Figures & tables

Appendix figures & tables24 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. MasDrift: Benchmarking Authorization Preservation Across Multi-Agent Architectures

    Aug 2, 2026Zhuoning Xu, Xiucheng Zhang, Hanjun Luo +3Model-Based Multi-Agent SystemsDelegation

  2. Governing What You Cannot Observe: Adaptive Runtime Governance for Autonomous AI Agents

    Apr 27, 2026German Marin, Jatin ChaudharyProbabilistic SafetyGovernance