Latent communication enables multi-agent systems to exchange information directly in internal representation space, reducing the token, computation, and latency overhead of text-based communication. To this end, lightweight trainable links are introduced to map the sender's representations into the receiver's input space. In this work, we show that even benign link training can increase harmful compliance relative to text-based communication while the underlying safety-aligned agents remain unchanged. An attacker can amplify this effect by optimizing the links on harmful query--response pairs or poisoning otherwise benign training data. We further develop a reinforcement-learning attack that rewards harmful compliance alongside benign task performance without requiring harmful target responses. Across three communication topologies and four safety benchmarks, this attack raises the mean harmful-compliance score from 27.9 with benignly trained links to 76.9. Compared with direct supervised optimization, it also achieves higher average accuracy on two benign utility benchmarks. Adapting the rewards toward safer behavior also enables repair of compromised links, substantially reducing harmful compliance across all evaluated attacks without updating the agents. Overall, our results show that safety alignment requires considering the multi-agent system as a whole. Code: https://github.com/Muhammad-Huzaifaa/latent-safety
Figures & tables
Figure 1: Latent communication topologies. Frozen agents communicate through trainable links fθ , which map sender hidden states hs to soft tokens z in the receiver’s embedding space. Each agent also receives the query q (not shown for clarity).
Safety (ASR %, ↓ better)
Utility (%, ↑ better)
Efficiency ( ↓ better)
Topology
Comm.
HB
SR
JBB
AB
Avg.
Δ
MATH500
GPQA-D
Tokens
Time (s)
2-Agent
Text
4.0
8.6
4.0
1.0
4.4
51.5
33.7
729
2.65
Latent
39.3
36.9
36.6
11.5
31.1
+26.7
65.3
35.7
511
1.96
Sequential
Text
6.0
6.4
0.0
0.2
3.2
58.3
26.6
1042
2.30
Latent
38.8
35.0
26.7
26.1
31.7
+28.5
59.5
30.2
501
1.84
Mixture
Text
19.4
18.8
9.9
1.7
12.5
80.6
34.7
1123
8.44
Table 1: Text-based vs. latent communication. Harmful compliance, benign utility, and inference cost for text-based and benignly trained latent communication across three topologies.
Table 2: Supervised and data-poisoning attacks. Harmful compliance and benign utility after direct supervised optimization and data poisoning across three communication topologies.
Figure 3: Attacking latent communication. Links are optimized using (a) harmful targets, directly or through data poisoning, or (b) response-level rewards. All agents remain frozen.
Harmful compliance (%, ↑ less safe)
Benign utility (%, ↑ better)
Topology
Condition
HB
SR
JBB
AB
Avg.
MATH500
GPQA-D
2-Agent
Clean
39.3
36.9
36.6
11.5
31.1
65.3
35.7
RL attack
78.6
74.2
86.1
86.4
81.3
63.7
25.1
Sequential
Clean
38.8
35.0
26.7
26.1
31.7
59.5
30.2
RL attack
61.7
65.6
50.5
36.7
53.6
68.7
28.1
Mixture
Clean
29.4
11.8
27.7
14.4
20.8
76.8
43.2
Table 3: Reward-guided attack on latent communication. Harmful compliance and benign utility after reward-guided link optimization across three communication topologies.
Figure 4: Attacking single links.
Harmful compliance (ASR %, ↓ better)
Benign utility (%, ↑ better)
Topology
Setting
HB
SR
JBB
AB
Avg.
Δ
MATH500
GPQA-D
Clean (no attack)
39.3
36.9
36.6
11.5
31.1
65.3
35.7
Poisoning
59.2
54.5
48.5
34.7
49.2
63.9
31.7
↪ repair
6.5
12.4
4.0
0.4
5.8
−43.4
64.5
25.6
Supervised
80.6
78.0
76.2
78.5
78.3
18.4
30.2
↪ repair
1.0
3.5
1.0
0.2
1.4
−76.9
56.3
29.1
Table 4: Reward-guided repair. Harmful compliance and benign utility after repairing the communication links. Δ denotes the change in mean harmful compliance relative to the attacked system.
Table 5: Multi-agent configurations used in our experiments. Only the latent communication links are trainable; all LLM parameters remain frozen.
Dataset
Role
Examples
Supervision
Sequential-Math ( Zou et al., 2026a )
Clean link training
1,904
Benign task supervision
PKU-SafeRLHF ( Ji et al., 2025 )
Supervised attack
3,000
Harmful query–response pairs
PKU-SafeRLHF ( Ji et al., 2025 )
Data poisoning
212
Harmful query–response pairs
PKU-SafeRLHF ( Ji et al., 2025 )
Target-free RL attack
2,697
Harmful queries only
Sequential-Math ( Zou et al., 2026a )
RL utility preservation
1,365
Benign question–answer pairs
HarmBench ( Mazeika et al., 2024 )
Safety evaluation
200
Queries only
Appendix
Table 6: Datasets used for link training, attacks, and evaluation. Supervised and poisoning attacks use harmful query–response pairs, whereas the target-free RL attack uses harmful queries only.
Harmful compliance (%, ↑ less safe)
Mean
Topology
Optimization
HB
SR
JBB
AB
2-Agent
Supervised
82.1
79.3
69.3
80.2
77.7
Reward-guided RL
76.6
85.4
77.2
79.3
79.6
Sequential
Supervised
66.7
70.4
63.4
73.1
68.4
Reward-guided RL
64.2
58.9
57.4
57.8
59.6
Mixture
Supervised
78.6
82.8
81.2
77.9
80.1
Appendix
Table 7: Supervised vs. reward-guided RL optimization. Starting from the same untrained latent-link initialization, both objectives successfully optimize the communication links toward harmful compliance. Reward-guided RL achieves higher mean harmful compliance in two of the three topologies.
Harmful compliance (%, ↓ better)
Utility (%, ↑ better)
Topology
Training
HB
SR
JBB
AB
Avg.
MATH500
2-Agent
Supervised
39.3
36.9
36.6
11.5
31.1
65.3
RL
10.0
15.9
5.0
0.6
7.9
66.9
Sequential
Supervised
38.8
35.0
26.7
26.1
31.7
59.5
RL
27.9
31.8
14.9
4.0
19.6
68.7
Mixture
Supervised
29.4
11.8
27.7
14.4
20.8
76.8
Appendix
Table 8: Supervised vs. RL link training on benign data. Both methods start from the same untrained latent-link initialization and use only benign mathematics data. RL consistently yields lower harmful compliance while matching or improving benign utility.
Sequential
Mixture
Poison rate
HB
SR
JBB
AB
Mean
Δ
HB
SR
JBB
AB
Mean
Δ
0% (clean)
39.0
34.8
27.0
26.2
31.7
—
29.4
11.8
27.7
14.4
20.8
—
10%
55.0
60.4
52.0
56.9
56.1
+24.4
72.0
72.2
68.0
67.3
69.9
+49.1
20%
64.0
69.0
71.0
72.5
69.1
+37.4
78.0
79.2
82.0
74.4
78.4
+57.6
30%
58.0
64.9
63.0
57.5
60.8
+29.1
79.0
74.1
80.0
72.5
76.4
+55.6
40%
53.0
63.9
57.0
54.0
57.0
+25.3
79.0
70.3
78.0
68.1
73.8
+53.0
Appendix
Table 9: Effect of poisoning rate on harmful compliance. Attack success increases sharply with limited poisoning and peaks around 20% in both topologies, after which additional contamination provides diminishing gains.
Harmful compliance (%, ↑ less safe)
Mean
Topology
Attacked link
HB
SR
JBB
AB
Sequential
Clean (no attack)
38.8
35.0
26.7
26.1
31.7
Planner → Refiner
61.7
64.6
57.4
65.1
62.2
Refiner → Solver
73.1
72.9
65.3
68.5
70.0
All links
68.2
69.7
61.4
68.5
67.0
Mixture
Clean (no attack)
29.4
11.8
27.7
14.4
20.8
Appendix
Table 10: Attack surface across individual latent links. Attacking a single communication link is sufficient to induce substantial harmful compliance, and joint optimization of all links is not consistently stronger than targeting the most vulnerable link.
Starting link
After RL
Topology
RL objective
ASR
MATH500
ASR
MATH500
Δ MATH
Degraded initialization
2-Agent
harm + utility
78.3
18.4
95.6
43.4
+25.0
Sequential
harm + utility
67.0
25.7
89.6
64.2
+38.5
Mixture
harm + utility
81.6
47.9
92.1
65.6
+17.7
Clean initialization
Appendix
Table 11: Effect of the utility reward. Utility-guided RL strongly recovers degraded links, while its benefit from a clean initialization is topology-dependent. Δ MATH is relative to the starting link.
Diagnostic
Condition / Intervention
Result
Forced refusal prefix
Receiver only
ASR: 9.0→1.5
Untrained link
ASR: 8.5→1.0
Benign-trained link
ASR: 31.1→0.0
Prefix-length intervention
Force I
69% of link-induced effect recovered
Force No
88% recovered
Force Sorry
95% recovered
Appendix
Table 12: Mechanistic diagnostics of benign supervised link training. The trained link primarily shifts the receiver away from refusal at the start of generation, while the underlying refusal behavior remains recoverable.
Figure 5: Qualitative examples in the simple two-agent communication setting. We compare receiver responses under the clean latent link, three link attacks, and the reward-guided defense. The examples illustrate that benign link training can already weaken refusal behavior, while adversarial link optimization further increases harmful compliance. The defense restores refusal behavior.
Figure 6: Qualitative examples in the simple two-agent communication setting.
Figure 7: Qualitative examples in the latent sequential communication setting.
Figure 8: Qualitative examples in the latent sequential communication setting.
Figure 9: Qualitative examples in the mixture style communication setting.
Figure 10: Qualitative examples in the mixture style communication setting.
Multi-agent systems rely on communication for information sharing and action coordination, which exposes a vulnerability to attacks. We investigate single-victim communication perturbation attacks against Multi-Agent Reinforcement Learning-trained systems and propose methods that use gradient information from the Jacobian to identify which messages, agent, and timesteps are most susceptible to attack and have the greatest impact on the system. We enhance these methods with two proposed adversarial loss functions that trade-off attack success for attack impact which also create more effective perturbations. We empirically demonstrate the effectiveness of our methods against two different multi-agent communication methods in navigation, PredatorPrey, and TrafficJunction environments. Our results show that our novel message selection method achieves a similar or greater impact than random message selection across almost all tested scenarios. Our victim selection, message selection, tempo, and loss functions improve attack effectiveness in half of the thirty scenarios we tested.
Latent-based multi-agent systems replace parts of explicit inter-agent communication with hidden representations, offering a new direction for efficient and flexible agent collaboration. However, moving coordination into latent space may also move attacks beyond the reach of visible-text inspection. In this paper, we study whether latent states can carry attack-associated information that remains effective during clean executions. To examine this question, we introduce a latent attack framework that reactivates attack-induced effects through latent interventions without reusing adversarial text. Extensive experiments show that the resulting latent-only attacks can substantially degrade task performance in clean executions, especially when applied to inter-agent KV-cache handoffs rather than local hidden states. Further control analyses indicate that this degradation cannot be reduced to arbitrary perturbations or invalid generation. Overall, our findings suggest that latent-based collaboration does not remove attack risk. It shifts part of the risk into less observable execution states, calling for safeguards beyond visible-text inspection.
Chenxi Wang, Ruiyang Huang, Jiayan Sun +2
Southeast University, Nanjing, China · Peking University, Beijing, China
Multi-agent systems built on large language models (LLMs) have become a prevailing paradigm for tackling complex reasoning, planning, and tool-use tasks. The dominant communication protocol in such systems is natural language: agents exchange messages token-by-token, verbalising their internal reasoning so that peers can read, verify, and respond. While convenient and interpretable, this protocol suffers from three structural drawbacks -- high inference cost, irreversible information loss during discretization, and ambiguity/redundancy of natural language. A growing body of work therefore explores an alternative protocol -- latent communication -- in which agents exchange continuous representations (embeddings, hidden states, or KV-caches) directly, bypassing the bottleneck of text generation. This paper presents a unified framework for organising the rapidly expanding literature on latent communication. We analyse existing methods along three orthogonal axes: (1) WHAT information is communicated (Embeddings, Hidden States, KV-Caches, or other continuous state); (2) WHICH sender-receiver alignment is used (latent-space alignment and layer alignment); and (3) HOW the communicated information is fused into the receiver (concatenation, prepending, mathematical operations, cross-attention, or cache restoration). Under this 3-axis framework, we systematically categorise eighteen representative methods proposed between 2024 and 2026, identify five major design patterns, and surface a set of open challenges -- including cross-architecture alignment, security of latent channels, compression for edge deployment, and the relationship between latent communication and latent chain-of-thought. We hope that this framework both lowers the barrier to entry for new researchers and provides a vocabulary for comparing future work.