Driven by the rapid advancement of large language models (LLMs), LLM-based multi-agent systems (MAS) have emerged as a powerful paradigm for collaborative reasoning over complex tasks. A key design element of MAS is the communication topology, which governs information flow among agents and often encodes proprietary knowledge about the system architecture. However, recent work has shown that such topologies can be inferred even in black-box settings by exploiting semantic dependencies in observable reasoning traces, posing significant risks of intellectual property leakage and exposure of system vulnerabilities. To address this threat, we propose MIRAGE, a topology-concealment framework that preserves the genuine communication topology for task execution while shaping adversary-facing semantic evidence toward a carefully constructed phantom topology. Specifically, MIRAGE operates in three stages: (1) phantom topology synthesis, (2) semantic edge realization, and (3) protected MAS execution. It constructs a phantom topology structurally distinct from the genuine one, materializes phantom edges as plausible semantic dependencies, and suppresses source-specific cues that could reveal genuine edges absent from the phantom topology. Extensive experiments across three topology optimization frameworks and four benchmark datasets demonstrate that MIRAGE substantially reduces the effectiveness of topology inference attacks while largely preserving the task utility of the protected MAS.
Figures & tables
Figure 1: Comparison of topology inference attack and Mirage defense. (a) The topology inference attack against LLM-based MAS exploits semantic dependencies in observable reasoning traces to infer the genuine communication topology G . (b) Mirage constructs a phantom topology G′ to mislead topology inference while preserving G for actual task execution. This design decouples topology exposure from task execution, effectively concealing the genuine communication topology.
Figure 2: Overview of Mirage . Mirage consists of three stages: ① phantom topology synthesis, ② semantic edge realization, and ③ protected MAS execution. Stage I constructs a phantom topology G′ that deviates from the genuine topology G . Stage II realizes G′ by shaping adversary-facing semantic dependencies. Stage III preserves G for task execution while exposing dependency evidence aligned with G′ , thereby concealing the genuine topology with minimal impact on task utility.
Figure 3: Comparison of topology inference performance using G-Designer before defense ( No Defense ) and with Mirage ( Ours ) across four benchmark datasets in terms of AUC, ACC, and F1.
Figure 4: Comparison of topology inference performance using AGP before defense ( No Defense ) and with Mirage ( Ours ) across four benchmark datasets in terms of AUC, ACC, and F1.
Figure 5: Comparison of topology inference performance using ARG-Designer before defense ( No Defense ) and with Mirage ( Ours ) across four benchmark datasets in terms of AUC, ACC, and F1.
Method
MMLU
GSM8K
SVAMP
HumanEval
AUC
ACC
F1
AUC
ACC
F1
AUC
ACC
F1
AUC
ACC
F1
Instruction
0.79
0.59
0.62
0.81
0.73
0.64
0.72
0.74
0.69
0.78
0.62
0.59
Delimiters
0.76
0.61
0.64
0.80
0.76
0.67
0.70
0.68
0.71
0.76
0.65
0.56
Mirage (Ours)
0.67
0.50
0.49
0.64
0.51
0.42
0.64
0.48
0.39
0.68
0.44
0.52
Table 1: Comparison of topology inference performance under different defense methods using G-Designer across four benchmark datasets in terms of AUC, ACC, and F1.
Figure 6: Overall ROC curves of PPL-based attack detection across four benchmark datasets, with the corresponding detection AUC reported and random guessing shown as a reference.
Figure 7: Comparison of task utility before defense ( Original ) and with Mirage ( Ours ) across three topology optimization strategies and four benchmark datasets, measured by task accuracy.
Figure 8: Ablation and parameter analyses of Mirage under different experimental settings. (a) Effects of removing key defense components. (b)–(d) Effects of the structural deviation threshold ρ0 , number of paraphrase candidates M , and dependency suppression weight λ , respectively.
Method
MMLU
GSM8K
SVAMP
HumanEval
Nˉ
Eˉ
Nˉ
Eˉ
Nˉ
Eˉ
Nˉ
Eˉ
G-Designer
7.00
8.99
5.00
8.19
5.00
8.15
6.00
11.38
AGP
6.00
10.87
5.00
8.45
5.00
8.41
6.00
11.54
ARG-Designer
5.42
7.84
3.07
3.14
3.05
3.10
4.24
5.49
Table 2: Statistics of communication topologies generated by three topology optimization methods across four datasets. Nˉ and Eˉ denote the average numbers of agents and edges, respectively.
Figure 9: Visualization of communication topologies under G-Designer, AGP, and ARG-Designer, comparing the ground-truth topology ( Ground-truth ) with those inferred by CIA before defense ( No Defense ) and with Mirage ( Ours ) across diverse topology structures.
Large language model (LLM)-powered multi-agent systems (MAS) enable agents to communicate and share information, achieving strong performance on complex tasks. However, this communication also creates an attack surface where malicious agents can propagate misinformation and manipulate group decisions, undermining MAS safety. Existing embedding-based defenses aim to detect and prune suspicious agents, but their effectiveness depends on a clear separation between the text embeddings of malicious and benign messages. Attackers can circumvent such defenses by crafting messages whose embeddings lie close to benign ones. We analyze this failure mode theoretically and validate it empirically with three attacks, Slow Drift, Benign Wrapper, and Chaos Seeding. Our analysis further reveals a fundamental limitation of embedding-based defenses: because they rely solely on the text embeddings, they ignore token-level confidence signals such as logits, which can remain informative when embeddings are not distinguishable under attack. We propose using confidence scores to prune or down-weight messages during MAS communication. Experiments show improved robustness across models, datasets, and communication topologies. Moreover, we find that the effectiveness of confidence signals decays over communication rounds, highlighting the importance of early intervention. This insights can inform and inspire future work on MAS attacks and defenses.
The performance of large language model (LLM)-based multi-agent systems (MAS) largely depends on effective communication topologies. Existing topology generation methods, however, typically learn communication topologies through black-box optimization driven solely by task-level rewards. While effective, such optimization provides little insight into why particular communication edges are selected, making it difficult to identify the critical communication subgraphs responsible for successful collaboration. To address this limitation, we propose E2-Explainer, a model-agnostic framework for providing interpretable explanations of communication topologies produced by arbitrary topology generators. Specifically, we formulate topology explanation as a causal attribution problem that identifies compact communication subgraphs supported by edge-level evidence of task preservation. We obtain this evidence with a Granger-style objective that measures how masking each communication channel changes the task outcome and the stability of the final response. The resulting budgeted subgraphs are then distilled into an amortized explainer, enabling efficient post-hoc explanation without repeated edge-level evaluations at deployment. Extensive experiments on multiple reasoning and coding benchmarks demonstrate that E2-Explainer identifies critical communication subgraphs that preserve successful collaboration. These subgraphs can also be executed directly to prune redundant communication edges, substantially reducing communication costs while maintaining competitive task performance.
Junzhi Li, Peng He, Qirui Ji +3
University of Chinese Academy of Sciences · National Key Laboratory of Space Integrated Information System, Institute of Software, Chinese Academy of Sciences · Beijing University of Posts and Telecommunications
LLM-based multi-agent systems (LLM-MAS) have become a promising paradigm for solving complex tasks through role specialization, tool use, memory, and collaborative reasoning. However, these interactions create new security risks that malicious instructions injected through messages, tools, or memories can propagate across agents and rounds, causing system-level compromise. Existing defenses largely rely on local filtering or graph-based anomaly detection, but they often fail to trace fine-grained propagation paths or remediate contaminated states without disrupting benign collaboration. We propose PropGuard, a propagation-aware framework for safeguarding LLM-MAS. PropGuard constructs a dual-view spatio-temporal graph that combines response-centric risk estimation with full-state evidence preservation. Guided by these risk priors, a GE-GRPO trained inspector sequentially explores the full-state graph to recover compact suspicious propagation subgraphs. PropGuard then verifies harmful propagation through subgraph-aware diagnosis and applies source-guided remediation to correct upstream contamination and replay affected downstream interactions. Experiments across four communication architectures and five attack settings demonstrate that PropGuard consistently lowers attack success while maintaining high task-level defense success, achieving a favorable effectiveness--efficiency trade-off.
Bingyu Yan, Xiaoming Zhang, Jinyu Hou +4
Beihang University · Beijing University of Posts and Telecommunications · City University of Hong Kong