Latent communication enables multi-agent systems to exchange information directly in internal representation space, reducing the token, computation, and latency overhead of text-based communication. To this end, lightweight trainable links are introduced to map the sender's representations into the receiver's input space. In this work, we show that even benign link training can increase harmful compliance relative to text-based communication while the underlying safety-aligned agents remain unchanged. An attacker can amplify this effect by optimizing the links on harmful query--response pairs or poisoning otherwise benign training data. We further develop a reinforcement-learning attack that rewards harmful compliance alongside benign task performance without requiring harmful target responses. Across three communication topologies and four safety benchmarks, this attack raises the mean harmful-compliance score from 27.9 with benignly trained links to 76.9. Compared with direct supervised optimization, it also achieves higher average accuracy on two benign utility benchmarks. Adapting the rewards toward safer behavior also enables repair of compromised links, substantially reducing harmful compliance across all evaluated attacks without updating the agents. Overall, our results show that safety alignment requires considering the multi-agent system as a whole. Code: https://github.com/Muhammad-Huzaifaa/latent-safety
Figures & tables
Figure 1: Latent communication topologies. Frozen agents communicate through trainable links fθ , which map sender hidden states hs to soft tokens z in the receiver’s embedding space. Each agent also receives the query q (not shown for clarity).
Safety (ASR %, ↓ better)
Utility (%, ↑ better)
Efficiency ( ↓ better)
Topology
Comm.
HB
SR
JBB
AB
Avg.
Δ
MATH500
GPQA-D
Tokens
Time (s)
2-Agent
Text
4.0
8.6
4.0
1.0
4.4
51.5
33.7
729
2.65
Latent
39.3
36.9
36.6
11.5
31.1
+26.7
65.3
35.7
511
1.96
Sequential
Text
6.0
6.4
0.0
0.2
3.2
58.3
26.6
1042
2.30
Latent
38.8
35.0
26.7
26.1
31.7
+28.5
59.5
30.2
501
1.84
Mixture
Text
19.4
18.8
9.9
1.7
12.5
80.6
34.7
1123
8.44
Table 1: Text-based vs. latent communication. Harmful compliance, benign utility, and inference cost for text-based and benignly trained latent communication across three topologies.
Table 2: Supervised and data-poisoning attacks. Harmful compliance and benign utility after direct supervised optimization and data poisoning across three communication topologies.
Figure 3: Attacking latent communication. Links are optimized using (a) harmful targets, directly or through data poisoning, or (b) response-level rewards. All agents remain frozen.
Harmful compliance (%, ↑ less safe)
Benign utility (%, ↑ better)
Topology
Condition
HB
SR
JBB
AB
Avg.
MATH500
GPQA-D
2-Agent
Clean
39.3
36.9
36.6
11.5
31.1
65.3
35.7
RL attack
78.6
74.2
86.1
86.4
81.3
63.7
25.1
Sequential
Clean
38.8
35.0
26.7
26.1
31.7
59.5
30.2
RL attack
61.7
65.6
50.5
36.7
53.6
68.7
28.1
Mixture
Clean
29.4
11.8
27.7
14.4
20.8
76.8
43.2
Table 3: Reward-guided attack on latent communication. Harmful compliance and benign utility after reward-guided link optimization across three communication topologies.
Figure 4: Attacking single links.
Harmful compliance (ASR %, ↓ better)
Benign utility (%, ↑ better)
Topology
Setting
HB
SR
JBB
AB
Avg.
Δ
MATH500
GPQA-D
Clean (no attack)
39.3
36.9
36.6
11.5
31.1
65.3
35.7
Poisoning
59.2
54.5
48.5
34.7
49.2
63.9
31.7
↪ repair
6.5
12.4
4.0
0.4
5.8
−43.4
64.5
25.6
Supervised
80.6
78.0
76.2
78.5
78.3
18.4
30.2
↪ repair
1.0
3.5
1.0
0.2
1.4
−76.9
56.3
29.1
Table 4: Reward-guided repair. Harmful compliance and benign utility after repairing the communication links. Δ denotes the change in mean harmful compliance relative to the attacked system.
Table 5: Multi-agent configurations used in our experiments. Only the latent communication links are trainable; all LLM parameters remain frozen.
Dataset
Role
Examples
Supervision
Sequential-Math ( Zou et al., 2026a )
Clean link training
1,904
Benign task supervision
PKU-SafeRLHF ( Ji et al., 2025 )
Supervised attack
3,000
Harmful query–response pairs
PKU-SafeRLHF ( Ji et al., 2025 )
Data poisoning
212
Harmful query–response pairs
PKU-SafeRLHF ( Ji et al., 2025 )
Target-free RL attack
2,697
Harmful queries only
Sequential-Math ( Zou et al., 2026a )
RL utility preservation
1,365
Benign question–answer pairs
HarmBench ( Mazeika et al., 2024 )
Safety evaluation
200
Queries only
Appendix
Table 6: Datasets used for link training, attacks, and evaluation. Supervised and poisoning attacks use harmful query–response pairs, whereas the target-free RL attack uses harmful queries only.
Harmful compliance (%, ↑ less safe)
Mean
Topology
Optimization
HB
SR
JBB
AB
2-Agent
Supervised
82.1
79.3
69.3
80.2
77.7
Reward-guided RL
76.6
85.4
77.2
79.3
79.6
Sequential
Supervised
66.7
70.4
63.4
73.1
68.4
Reward-guided RL
64.2
58.9
57.4
57.8
59.6
Mixture
Supervised
78.6
82.8
81.2
77.9
80.1
Appendix
Table 7: Supervised vs. reward-guided RL optimization. Starting from the same untrained latent-link initialization, both objectives successfully optimize the communication links toward harmful compliance. Reward-guided RL achieves higher mean harmful compliance in two of the three topologies.
Harmful compliance (%, ↓ better)
Utility (%, ↑ better)
Topology
Training
HB
SR
JBB
AB
Avg.
MATH500
2-Agent
Supervised
39.3
36.9
36.6
11.5
31.1
65.3
RL
10.0
15.9
5.0
0.6
7.9
66.9
Sequential
Supervised
38.8
35.0
26.7
26.1
31.7
59.5
RL
27.9
31.8
14.9
4.0
19.6
68.7
Mixture
Supervised
29.4
11.8
27.7
14.4
20.8
76.8
Appendix
Table 8: Supervised vs. RL link training on benign data. Both methods start from the same untrained latent-link initialization and use only benign mathematics data. RL consistently yields lower harmful compliance while matching or improving benign utility.
Sequential
Mixture
Poison rate
HB
SR
JBB
AB
Mean
Δ
HB
SR
JBB
AB
Mean
Δ
0% (clean)
39.0
34.8
27.0
26.2
31.7
—
29.4
11.8
27.7
14.4
20.8
—
10%
55.0
60.4
52.0
56.9
56.1
+24.4
72.0
72.2
68.0
67.3
69.9
+49.1
20%
64.0
69.0
71.0
72.5
69.1
+37.4
78.0
79.2
82.0
74.4
78.4
+57.6
30%
58.0
64.9
63.0
57.5
60.8
+29.1
79.0
74.1
80.0
72.5
76.4
+55.6
40%
53.0
63.9
57.0
54.0
57.0
+25.3
79.0
70.3
78.0
68.1
73.8
+53.0
Appendix
Table 9: Effect of poisoning rate on harmful compliance. Attack success increases sharply with limited poisoning and peaks around 20% in both topologies, after which additional contamination provides diminishing gains.
Harmful compliance (%, ↑ less safe)
Mean
Topology
Attacked link
HB
SR
JBB
AB
Sequential
Clean (no attack)
38.8
35.0
26.7
26.1
31.7
Planner → Refiner
61.7
64.6
57.4
65.1
62.2
Refiner → Solver
73.1
72.9
65.3
68.5
70.0
All links
68.2
69.7
61.4
68.5
67.0
Mixture
Clean (no attack)
29.4
11.8
27.7
14.4
20.8
Appendix
Table 10: Attack surface across individual latent links. Attacking a single communication link is sufficient to induce substantial harmful compliance, and joint optimization of all links is not consistently stronger than targeting the most vulnerable link.
Starting link
After RL
Topology
RL objective
ASR
MATH500
ASR
MATH500
Δ MATH
Degraded initialization
2-Agent
harm + utility
78.3
18.4
95.6
43.4
+25.0
Sequential
harm + utility
67.0
25.7
89.6
64.2
+38.5
Mixture
harm + utility
81.6
47.9
92.1
65.6
+17.7
Clean initialization
Appendix
Table 11: Effect of the utility reward. Utility-guided RL strongly recovers degraded links, while its benefit from a clean initialization is topology-dependent. Δ MATH is relative to the starting link.
Diagnostic
Condition / Intervention
Result
Forced refusal prefix
Receiver only
ASR: 9.0→1.5
Untrained link
ASR: 8.5→1.0
Benign-trained link
ASR: 31.1→0.0
Prefix-length intervention
Force I
69% of link-induced effect recovered
Force No
88% recovered
Force Sorry
95% recovered
Appendix
Table 12: Mechanistic diagnostics of benign supervised link training. The trained link primarily shifts the receiver away from refusal at the start of generation, while the underlying refusal behavior remains recoverable.
Figure 5: Qualitative examples in the simple two-agent communication setting. We compare receiver responses under the clean latent link, three link attacks, and the reward-guided defense. The examples illustrate that benign link training can already weaken refusal behavior, while adversarial link optimization further increases harmful compliance. The defense restores refusal behavior.
Figure 6: Qualitative examples in the simple two-agent communication setting.
Figure 7: Qualitative examples in the latent sequential communication setting.
Figure 8: Qualitative examples in the latent sequential communication setting.
Figure 9: Qualitative examples in the mixture style communication setting.
Figure 10: Qualitative examples in the mixture style communication setting.