Multi-agent LLM systems coordinate task execution through exchanges of information among agents. When coordination breaks down, similar symptoms in execution traces can reflect different problems in how information is passed, used, or verified. Communication topology captures how agents exchange information and provides structural cues for distinguishing coordination failure modes. Using these cues for diagnosis requires establishing how topology relates to failure patterns and recovering the relevant structure from execution traces that lack explicit topology labels. We analyze the relationship between communication topology and failure patterns and introduce MAScope, a two-stage framework for topology-conditioned diagnosis. Its Trace Structural Extractor TSE recovers communication topology from heterogeneous execution traces by grounding an interaction graph in message evidence. The Topology-Conditioned Judge TC-Judge then classifies failures using the trace, predicted topology, an empirical failure prior estimated from separate labeled traces, and a short description of topology-specific failure patterns. Under a fixed orchestration structure, the recovered topology can be reused across executions. Experimental results show a statistically significant association between communication topology and failure type, with χ2=409.9 and p=1.2×10−70. On the \num{851} MAST-clean traces, ground-truth topology context raises gpt-mini's Macro-F1 from 0.173 to 0.350. With predicted topology, the pipeline achieves 0.346, approaching the trace-only gpt-5.4 baseline of 0.372. For \num{1000} traces under a fixed orchestration structure, the projected pipeline cost, including one topology extraction, is approximately 6% of repeated gpt-5.4 diagnosis cost. These results show that topology-conditioned context improves failure diagnosis and supports lower-cost deployment.
Figures & tables
Figure 1. Shared message pool context guides checks of feedback use and task verification. In this trace, the reviewer’s feedback was not acted upon and the requested verification was not performed. A shared-message-pool execution in which a reviewer requests interrupted-transfer handling and verification, but the tester instead checks an unrelated square function and the next review accepts it. The topology view highlights the unaddressed feedback and missing task verification.
Corpus
# traces
Lay.
Cen.
Dec.
SMP
MAST-clean
851
256
165
0
430
LG-Ctrl
408
102
102
102
102
Inject-Mix
961
170
284
507
0
Table 1. MAScope-Bench composition. Lay., Cen., Dec., and SMP denote Layered, Centralized, Decentralized, and Shared Message Pool, based on each declared orchestration graph.
Figure 2. Topology–failure association on MAST-clean ( 851 traces). Cells show standardised Pearson residuals for the 4×14 contingency table, with ∣r∣≥2 annotated. The independence test gives χ2=409.9 , dof=39 , and p<10−70 . A four-by-fourteen heat map of standardized Pearson residuals between communication topology and MAST failure codes. Each topology row has a different cluster of positive residuals, with Layered associated with information loss, Centralized with verification misalignment, Decentralized with coordination stall, and Shared Message Pool with semantic drift.
Figure 3. MAScope pipeline. TSE converts a heterogeneous trace into an evidence-grounded relation graph and topology label τ^ . TC-Judge combines the trace, predicted topology, frozen topology prior, and failure signature to output a MAST-14 diagnosis. A two-stage pipeline. The first stage adapts and compacts a heterogeneous trace, extracts and validates evidence-grounded relations, and assigns a topology label. The second stage combines the original trace with the topology label, a topology-specific prior, and a failure signature to produce a fourteen-bit MAST diagnosis.
Family
Method
Setting
Macro-F1
Gain over matched M0
External protocol
MAST-Judge
gpt-mini
0.131
–
External protocol
MAST-Judge
gpt-5.4
0.363
–
Matched control
M0
gpt-mini
0.173
Reference
Matched control
M0
gpt-5.4
0.372
Reference
Ours
M4
gpt-mini with true topology
0.350
+0.177 and +102%
Ours
M4
gpt-5.4 with true topology
0.399
+0.027 and +7.3%
Table 2. Zero-shot diagnostic results on MAST-clean . No method uses its failure labels for fitting. The final column reports the gain from M0 to M4 on the same model.
Table 3. Baselines used in the main and cross-framework comparisons. All inputs are constructed from the same raw traces.
Figure 4. Controlled diagnostic gains from M0 to M4 on MAST-clean . Each comparison fixes the traces, prompt scaffolding, output space, and model. Both models improve, with gpt-mini gaining 102% . A grouped bar chart comparing MAST-Judge, the topology-free M-zero judge, and the topology-conditioned M-four judge on gpt-mini and gpt-5.4. M-four raises Macro-F1 from 0.173 to 0.350 for gpt-mini and from 0.372 to 0.399 for gpt-5.4.
Evaluation slice
N
C, W, A
Accuracy
MAST-clean by framework
ChatDev
256
211, 44, 1
82.4%
Magentic
165
160, 5, 0
97.0%
MetaGPT
430
427, 0, 3
99.3%
MAST-clean overall
851
798, 49, 4
93.8%
LG-Ctrl by construction topology
Table 4. Topology recovery by framework and controlled topology. C, W, and A denote correct, wrong, and abstained predictions.
Figure 5. Cross-framework scores on MAST-clean . Flat-GBM and Neural-Trace use strict LOFO retraining. M0 has no fitted component. M4∗ uses a full-corpus topology prior and is a descriptive framework slice rather than a strict LOFO result. Right labels report retention. A dumbbell chart comparing in-distribution and leave-one-framework-out Macro-F1. Flat-GBM falls from 0.490 to 0.226 and Neural-Trace from 0.296 to 0.127, while descriptive zero-shot slices for M-zero and M-four retain 63 percent and 82 percent respectively.
Figure 6. Cost comparison in USD. The upper panel reports measured gpt-5.4 cost and the normalized-token gpt-mini estimate on 851 traces without TSE . The lower panel adds one TSE call and shows deployment cost when the extracted topology is reused. Two cost charts. The upper bar chart compares 8.73 dollars for gpt-5.4 raw-trace diagnosis with 0.52 dollars for gpt-mini using topology on 851 traces. The lower log-scale line chart shows that reusing one extracted topology costs 0.63 dollars for one thousand gpt-mini diagnoses versus 10.26 dollars for gpt-5.4.