cs.AISep 24, 2026

MeshHeal: Two-Timescale Self-Healing for Gray Failures in Decentralized LLM Agent Networks

Authors: Keru Chen, Sen Lin, Yingbin Liang, Nathaniel D. Bastian, Shaofeng Zou

Organizations: School of ECEE Arizona State University · Computer Science Department University of Houston · Department of ECE The Ohio State University · Whiting School of Engineering Johns Hopkins University

Abstract

Decentralized LLM-based multi-agent systems coordinate through local interactions, but an agent can remain responsive while its task-solving quality persistently degrades. Such gray failures require protecting current tasks before sufficient evidence exists to alter future routing, while still allowing recovered agents to rejoin. We introduce MeshHeal, a fully decentralized self-healing framework that couples ability-matched peer review across two timescales. At the fast timescale, an adaptive hierarchy escalates uncertain or low-scoring outputs from repeated single-reviewer evaluation to committee deliberation and, when needed, correction before use. At the slow timescale, a task- and ability-conditioned peer-relative detector aggregates scores to distinguish persistent degradation from ordinary output variation, trigger mandatory committee review, and eventually exclude degraded agents from ordinary routing; recovery probes provide fresh evidence for reintegration. To faithfully evaluate routing, we introduce Model-Backed MAS Evaluation, which ties ability assignments to execution models, since prompt-based ability assignments alone can leave routing errors hidden. Across BBH, MATH, and MMLU-Pro, MeshHeal achieves 0.839 degraded-phase accuracy using 51k total model tokens per task, versus the strongest baseline Symphony's 0.807 accuracy using 115k per task. Under staggered degradation and recovery, MeshHeal isolates degraded agents, keeps them excluded from ordinary task execution until recovery, and returns them to normal routing.

Figures & tables

Appendix figures & tables21 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Robust Multi-Agent LLMs under Byzantine Faults

    May 9, 2026Haejoon Lee, Vincent-Daniel Yun, Dimitra Panagou +1Multi-Agent Large Language Model SystemsByzantine Attacks

  2. Who Broke the System? Failure Localization in LLM-Based Multi-Agent Systems

    Jul 8, 2026Yufei Xia, Anjun Gao, Yueyang Quan +2Model-Based Multi-Agent Systems

  3. Self-Evolving Multi-Agent Systems via Decentralized Memory

    May 21, 2026Guangya Hao, Yunbo Long, Zhuokai ZhaoMulti-Agent EvolutionNativemem