LLM-based multi-agent systems (MAS) increasingly use latent collaboration to avoid the information loss and repeated encoding-decoding overhead of natural-language communication. However, directly forwarding all sender latents makes the receiver-side context scale with both the number of agents and the reasoning length, increasing computation, memory usage, and collaboration latency. A natural solution is latent compression. But we find that cross-agent redundancy remains unresolved in existing latent compression approaches, which typically compress each sender independently and then concatenate the results. We propose LatCom, a cross-agent latent compression framework for efficient multi-agent latent collaboration. LatCom maps multiple sender latents into a fixed number of receiver-readable and task-relevant slots. Rather than reconstructing all sender hidden states, it optimizes the compressed latents for receiver-side task utility. LatCom trains the compressor in two stages: single-sender readability learning first establishes a latent interface interpretable by the frozen receiver, and multi-sender fusion learning then trains the compressor to fuse complementary evidence and remove redundancy across agents. Experiments on multiple benchmarks with Qwen3-4B show that LatCom achieves an average 2.46x inference speed-up over LatentMAS and reduces output token usage by 70.3% while maintaining comparable average accuracy.
Figures & tables
Figure 1: Independent per-agent compression preserves cross-agent redundancy, as t-SNE visualization shows overlap among concatenated compressed latents.
Figure 2: Overview of LatCom. Sender agents generate aligned latent thoughts, which LatCom aggregates and compresses into fixed-size receiver-readable slots. The frozen receiver uses these slots for final generation, while the compressor is trained through readability and fusion learning.
Tasks
Metrics
TextMAS
LatentMAS
InterLat
LatentMAS-H2O
LatentMAS-hidden
LatCom
Qwen3-4B
GSM8K
Acc.
89.40 ( ↑ 0.49%)
88.10 ( ↑ 1.98%)
85.82 ( ↑ 4.68%)
83.55 ( ↑ 7.53%)
86.58 ( ↑ 3.77%)
89.84
ARC-E
Acc.
96.93 ( ↑ 0.65%)
95.45 ( ↑ 2.21%)
95.09 ( ↑ 2.59%)
94.74 ( ↑ 2.98%)
95.71 ( ↑ 1.93%)
97.56
ARC-C
Acc.
91.66 ( ↑ 1.25%)
91.72 ( ↑ 1.19%)
90.78 ( ↑ 2.23%)
89.85 ( ↑ 3.29%)
90.44 ( ↑ 2.62%)
92.81
MedQA
Acc.
65.49 ( ↓ 1.25%)
66.33 ( ↓ 2.50%)
64.16 ( ↑ 0.79%)
62.00 ( ↑ 4.31%)
66.67 ( ↓ 3.00%)
64.67
MBPP+
Acc.
69.12 ( ↓ 2.40%)
66.67 ( ↑ 1.18%)
65.74 ( ↑ 2.62%)
64.81 ( ↑ 4.09%)
64.81 ( ↑ 4.09%)
67.46
Table 1: Main results of LatCom on 7 public benchmarks under the MAS setting. Values in parentheses report the relative accuracy change of LatCom over each baseline. ↑ and ↓ indicate higher and lower accuracy, respectively.
Figure 3: Inference speed-up across seven benchmarks.
Figure 4: Output token usage across seven benchmarks.
Figure 5: t-SNE visualization of latent communication
Figure 6: Receiver Entropy under Latent Conditioning
Figure 7: Ablation Study of Two-Stage Training
Stage
Setting
Acc.
Drop
Single-sender readability learning
Stage 1
Full Stage 1
80.06
–
Stage 1
w/o Lcon
74.53
↓ 5.53
Stage 1
w/o Lalign
68.41
↓ 11.65
Multi-source fusion learning
Stage 1 + Stage 2
Full training
89.84
–
Table 2: Ablation of training objectives on GSM8K. “Drop” is computed against the corresponding full setting within the same stage.
Figure 8: Sensitivity Analysis of Sender Number
Figure 9: Sensitivity to Compressed Slot Number
Appendix figures & tables7 assets
Supplementary material from the paper’s appendix.
Appendix
Dataset
Agent pair
Raw AvgCos
Centered AvgCos
Linear CKA
GSM8K
Math–Science
0.605
0.061
0.600
Math–Code
0.693
0.061
0.559
Science–Code
0.771
0.069
0.508
ARC-Easy
Math–Science
0.587
0.068
0.577
Math–Code
0.615
0.053
0.463
Science–Code
0.747
0.076
0.457
Appendix
Table 3: Pairwise similarity between latent trajectories from the math, science, and code agents.
Dataset
Math
Science
Code
Separate Sum
Joint Rank
Rank Compression
GSM8K
7.6
10.3
9.6
27.5
15.6
43.3%
ARC-Easy
7.3
9.8
9.1
26.2
14.8
43.5%
ARC-Challenge
8.0
10.8
10.0
28.8
16.3
43.4%
Appendix
Table 4: Effective-rank analysis of individual and joint agent trajectories.
Method
Slots
GSM8K Acc.
Joint Mean Pool
64
78.92
Single-source Continuation
64
86.47
LatCom
64
89.84
Appendix
Table 5: Controlled comparison of cross-agent fusion under greedy decoding.
Figure 10: Example of constructing evidence-structured sender inputs and a rationale-augmented target from a HotpotQA instance.
Category
Hyperparameter
Value
Optimization settings
Optimizer
Optimizer type
AdamW, β1=0.9 , β2=0.95 , ϵ=10−8
Optimizer
Stage 1 learning rate
2×10−5
Optimizer
Stage 2 learning rate
1×10−5
Optimizer
LR scheduler
Cosine decay
Optimizer
Warmup steps / ratio
5%
Appendix
Table 6: Optimization and loss hyperparameters for LatCom training.
Dataset
Category
Max. tokens
GSM8K
Math
2048
ARC-Easy
Commonsense sci.
2048
ARC-Challenge
Commonsense sci.
2048
MedQA
Medical QA
4096
MBPP-Plus
Code
4096
HumanEval-Plus
Code
4096
Appendix
Table 7: Evaluation benchmarks and maximum output lengths.
Category
Metric
Value
Interpretation
Global compression time
Runtime
GSM8K
0.095 s
Time for mapping U to M=Cϕ(U)
Runtime
ARC-Easy
0.096 s
Time for mapping U to M=Cϕ(U)
Runtime
ARC-Challenge
0.107 s
Time for mapping U to M=Cϕ(U)
Runtime
MedQA
0.126 s
Time for mapping U to M=Cϕ(U)
Runtime
MBPP-Plus
0.128 s
Time for mapping U to M=Cϕ(U)
Appendix
Table 8: Standalone overhead of the LatCom compressor. The model footprint is persistent, while CUDA and CPU deltas measure the additional peak memory introduced by the compression call itself.
Although large language model (LLM) based multi-agent systems (MAS) show their capability to solve complex tasks and achieve higher performance over single agent systems, they lead to huge computational overheads because of heavy communication between agents. Previous research has made efforts to train a sparse multi-agent graph or fine-tune a planner to orchestrate the workflow better. However, such extra training processes introduce computational costs and limit MAS to specific domains, therefore compromising their generalizability. In this paper, we propose CONCAT, a training-free multi-agent collaboration framework based on CONsensus and Confidence-driven Ad hoc Teaming to efficiently organize agent interactions. Specifically, agents are clustered based on their initial answers, and leaders of each cluster are selected based on the agents' confidence. Then, a heuristic function based on the Theory of Mind is designed to predict the collaboration benefits between every two leaders according to their answers and confidence. Finally, an ad hoc multi-agent network is organized after evicting a percentage of communications based on the predicted benefits. Experiments across three LLMs and three benchmarks show that CONCAT achieves up to 2.02x higher efficiency (accuracy/latency ratio) than LLM-Debate and outperforms training-aware methods such as AgentDropout, while reducing average latency by 50.1% on Qwen2.5-14B-Instruct, without any task-specific training.
Multi-agent systems built on large language models (LLMs) have become a prevailing paradigm for tackling complex reasoning, planning, and tool-use tasks. The dominant communication protocol in such systems is natural language: agents exchange messages token-by-token, verbalising their internal reasoning so that peers can read, verify, and respond. While convenient and interpretable, this protocol suffers from three structural drawbacks -- high inference cost, irreversible information loss during discretization, and ambiguity/redundancy of natural language. A growing body of work therefore explores an alternative protocol -- latent communication -- in which agents exchange continuous representations (embeddings, hidden states, or KV-caches) directly, bypassing the bottleneck of text generation. This paper presents a unified framework for organising the rapidly expanding literature on latent communication. We analyse existing methods along three orthogonal axes: (1) WHAT information is communicated (Embeddings, Hidden States, KV-Caches, or other continuous state); (2) WHICH sender-receiver alignment is used (latent-space alignment and layer alignment); and (3) HOW the communicated information is fused into the receiver (concatenation, prepending, mathematical operations, cross-attention, or cache restoration). Under this 3-axis framework, we systematically categorise eighteen representative methods proposed between 2024 and 2026, identify five major design patterns, and surface a set of open challenges -- including cross-architecture alignment, security of latent channels, compression for edge deployment, and the relationship between latent communication and latent chain-of-thought. We hope that this framework both lowers the barrier to entry for new researchers and provides a vocabulary for comparing future work.
Communication in Large Language Model (LLM)-based multi-agent systems is moving beyond discrete tokens to preserve richer context. Recent work such as LatentMAS enables agents to exchange latent messages through full key-value (KV) caches. However, full KV relay incurs high memory and communication cost. We adapt KV-cache eviction methods to this setting and introduce \textbf{Orthogonal BackFill (OBF)} to mitigate information loss from hard eviction. OBF injects a low-rank orthogonal residual from discarded KV states into the retained KV states. We evaluate OBF against full KV relay on nine benchmarks spanning mathematical reasoning, expert and commonsense QA, and coding. With only 9.9%-20.2% of the prompt KV states retained, H-OBF delivers between 97 and 120 of full KV relay's per-benchmark accuracy across the nine benchmarks. This suggests that more information does not necessarily lead to better communication; preserving the most useful information matters more. Our codebase is included in the supplementary material. Our codebase is publicly available on https://github.com/markli404/When-Less-Latent-Leads-to-Better-Relay.
Yiping Li, Zhiyu An, Wan Du
Department of Computer Science and Engineering University of California, Merced Merced, CA 95343