MeshHeal: Two-Timescale Self-Healing for Gray Failures in Decentralized LLM Agent Networks
Organizations: School of ECEE Arizona State University · Computer Science Department University of Houston · Department of ECE The Ohio State University · Whiting School of Engineering Johns Hopkins University
Abstract
Decentralized LLM-based multi-agent systems coordinate through local interactions, but an agent can remain responsive while its task-solving quality persistently degrades. Such gray failures require protecting current tasks before sufficient evidence exists to alter future routing, while still allowing recovered agents to rejoin. We introduce MeshHeal, a fully decentralized self-healing framework that couples ability-matched peer review across two timescales. At the fast timescale, an adaptive hierarchy escalates uncertain or low-scoring outputs from repeated single-reviewer evaluation to committee deliberation and, when needed, correction before use. At the slow timescale, a task- and ability-conditioned peer-relative detector aggregates scores to distinguish persistent degradation from ordinary output variation, trigger mandatory committee review, and eventually exclude degraded agents from ordinary routing; recovery probes provide fresh evidence for reintegration. To faithfully evaluate routing, we introduce Model-Backed MAS Evaluation, which ties ability assignments to execution models, since prompt-based ability assignments alone can leave routing errors hidden. Across BBH, MATH, and MMLU-Pro, MeshHeal achieves 0.839 degraded-phase accuracy using 51k total model tokens per task, versus the strongest baseline Symphony's 0.807 accuracy using 115k per task. Under staggered degradation and recovery, MeshHeal isolates degraded agents, keeps them excluded from ordinary task execution until recovery, and returns them to normal routing.
Figures & tables
| Model | Unassigned ability acc. | Assigned ability acc. | |
|---|---|---|---|
| DeepSeek-Chat | .029 | ||
| Llama-3.1-70B | .015 | ||
| GPT-4o-mini | .027 | ||
| GPT-OSS-120B | .046 | ||
| Qwen-2.5-7B | .058 |
| Method | H acc. | D acc. | Tokens |
|---|---|---|---|
| MeshHeal | |||
| Symphony | |||
| SAC | |||
| RAPS | |||
| AgentNet | |||
| AutoGen |
Appendix figures & tables21 assets
Supplementary material from the paper’s appendix.
Appendix
| Model | Condition | Unassigned ability acc. | Assigned ability acc. | Accuracy gap | Additional gap |
| DeepSeek-Chat | No statement | 0.812 | 0.800 | – | |
| Told unreliable | 0.775 | 0.792 | 0.017 | 0.029 | |
| Llama-3.1-70B | No statement | 0.781 | 0.783 | 0.002 | – |
| Told unreliable | 0.750 | 0.767 | 0.017 | 0.015 | |
| GPT-4o-mini | No statement | 0.744 | 0.708 | – | |
| Told unreliable | 0.675 | 0.667 | 0.027 |
| Agent | Declared abilities |
|---|---|
| A1 | reasoning, mathematical, knowledge |
| A2 | reasoning, mathematical, knowledge |
| A3 (degraded) | mathematical, knowledge, sequence, inference |
| A4 | language, sequence, spatial, inference |
| A5 (degraded) | reasoning, language, spatial |
| A6 | language, sequence, spatial, inference |
| Task type | Required abilities |
|---|---|
| causal_judgement | reasoning + inference |
| formal_fallacies | reasoning + inference |
| tracking_shuffled_objects_five_objects | reasoning + sequence |
| geometric_shapes | mathematical + spatial |
| object_counting | mathematical + spatial |
| date_understanding | mathematical + language |
| Agent | Declared abilities |
|---|---|
| A1 | symbolic_manipulation , numeric_computation , theorem_knowledge |
| A2 | symbolic_manipulation , numeric_computation , theorem_knowledge |
| A3 (degraded) | numeric_computation , theorem_knowledge , enumeration , multi_step_deduction |
| A4 | enumeration , multi_step_deduction , problem_translation , figure_reasoning |
| A5 (degraded) | symbolic_manipulation , problem_translation , figure_reasoning |
| A6 | enumeration , multi_step_deduction , problem_translation , figure_reasoning |
| Subject | Required abilities | Rationale |
|---|---|---|
| prealgebra | numeric_computation problem_translation | word-problem modeling + arithmetic |
| algebra | symbolic_manipulation multi_step_deduction | expression transformation + multi-step solving |
| geometry | numeric_computation figure_reasoning | geometric relations + numerical calculation |
| number_theory | symbolic_manipulation enumeration | algebraic constraints + systematic cases |
| counting_and_probability | symbolic_manipulation enumeration | counting expressions + complete case coverage |
| intermediate_algebra | symbolic_manipulation multi_step_deduction | deeper transformations + longer deduction |
| Agent | Declared abilities |
|---|---|
| A1 | humanities_knowledge , science_knowledge , quantitative_methods |
| A2 | humanities_knowledge , science_knowledge , quantitative_methods |
| A3 (degraded) | science_knowledge , quantitative_methods , scenario_inference , stepwise_analysis |
| A4 | scenario_inference , stepwise_analysis , text_comprehension , structural_reasoning |
| A5 (degraded) | humanities_knowledge , text_comprehension , structural_reasoning |
| A6 | scenario_inference , stepwise_analysis , text_comprehension , structural_reasoning |
| Domain | Required abilities | Rationale |
|---|---|---|
| health | science_knowledge text_comprehension | medical knowledge + comprehension of long questions |
| history | humanities_knowledge scenario_inference | historical facts + contextual inference |
| biology | science_knowledge structural_reasoning | biology knowledge + structural/process reasoning |
| psychology | humanities_knowledge scenario_inference | theory + scenario judgment |
| philosophy | humanities_knowledge scenario_inference | doctrine + argumentative stance |
| economics | quantitative_methods text_comprehension | interpretation + model/calculation |
| Dataset | BBH | MATH | MMLU-Pro | Overall |
|---|---|---|---|---|
| Exact pair match | 100% | 100% | 100% | 100% |
| Execution condition | Model |
|---|---|
| Declared ability, healthy agent | openai/gpt-oss-120b |
| Missing ability or degraded agent | meta-llama/llama-3.2-1b-instruct |
| Dataset | Scoring rule |
|---|---|
| BIG-Bench Hard | Exact answer extraction and matching |
| MATH | Mathematical equivalence with math_equal |
| MMLU-Pro | Extracted option letter |
| Label | Source | Implementation used in our framework |
|---|---|---|
| AgentNet | Yang et al. (2025b) | Native decentralized routing with adaptive edge weights and pruning. |
| RAPS | Li et al. (2026) | Peer reputation updates local links and routing avoids low-reputation peers. |
| AutoGen | Wu et al. (2023) | A mandatory GroupChatManager-style hub delegates work, receives worker outputs, and produces the final answer. |
| SAC | Lee et al. (2026) | Three teams produce answers, run two filtering and refinement rounds, and aggregate by majority. |
| Symphony | Guan et al. (2026) | GlobalLinUCB selects agents with the required abilities; three plans run in parallel; the final result is selected by voting and used to update routing statistics. |
| Parameter | Value |
|---|---|
| Committee reviewer slots | 2 |
| Committee deliberation rounds | 2 |
| Takeover correction candidates | 2 |
| Routing hop limit | 4 hops |
| Maximum agents per task | 6 |
| Maximum reviewed outputs per task | 3 |
| Agents | Sparse edges | Fraction of complete graph | Maximum ability distance |
| 96 | 228 | 5.00% | 3 |
| 288 | 705 | 1.71% | 3 |
| 576 | 1454 | 0.88% | 3 |
| Method | BBH | MATH | MMLU-Pro | Mean | Change |
|---|---|---|---|---|---|
| MeshHeal | |||||
| Symphony | |||||
| SAC | |||||
| RAPS | |||||
| AgentNet | |||||
| AutoGen |
| Method | H calls | D calls | H tokens (k) | D tokens (k) |
|---|---|---|---|---|
| MeshHeal | ||||
| Symphony | ||||
| SAC | ||||
| RAPS | ||||
| AgentNet | ||||
| AutoGen |
| Method | Healthy acc. | Degraded acc. | Change | Tokens/task |
|---|---|---|---|---|
| MeshHeal | ||||
| Symphony | ||||
| SAC | ||||
| RAPS | ||||
| AgentNet | ||||
| AutoGen |
| Method | Healthy acc. | Degraded acc. | Change | Tokens/task |
|---|---|---|---|---|
| MeshHeal | ||||
| Symphony | ||||
| SAC | ||||
| RAPS | ||||
| AgentNet | ||||
| AutoGen |
| Agent | True degraded interval | State sequence |
|---|---|---|
| A5 | 127–378 | first evidence at task 3; normal during warm-up 129 watched 132 isolated 422 watched 430 normal |
| A3 | 253–504 | first evidence at task 4; normal during warm-up 267 watched 268 isolated 548 watched 550 normal |
| Metric | Full method | Single reviewer | No Takeover |
|---|---|---|---|
| Healthy accuracy | |||
| Degraded-phase accuracy | |||
| Paired phase change |
| Pair A | Pair B | ||||
|---|---|---|---|---|---|
| Rule | Evidence | Missed | False | Missed | False |
| Absolute threshold | Own scores | 0/2 | 4/4 | 0/2 | 0/4 |
| CUSUM | Own scores | 0/2 | 2/4 | 1/2 | 0/4 |
| Relative peer comparison | Scores relative to peers | 0/2 | 0/4 | 0/2 | 0/4 |
| (a) Initial review | |
|---|---|
| Metric | Value |
| Mean score, healthy-agent outputs | |
| Mean score, degraded-agent outputs | |
| Recall on degraded-agent outputs | |
| False-alarm rate on healthy-agent outputs | |
| Committee escalation rate for degraded-agent outputs | |