Heterogeneous multi-agent systems combine models with different capabilities through a common communication interface. Exchanging internal states directly requires translating between model-specific representations and controlling intermediate computation. We introduce the Vision Wormhole, which repurposes the visual input interface of Vision-Language Models (VLMs) for continuous communication between frozen heterogeneous agents. A Universal Visual Codec encodes each sender's latent rollout into a fixed-size message, maps it through a shared reference space, and decodes received messages into the receiver's image-token span. Per-model codecs and affine reference maps form a hub-and-spoke architecture with O(N) components for N models. Each model learns its codec independently through self-distillation on anchor texts, and shared-anchor alignment enables reuse across communication partners. Across four VLM families, six team configurations, and nine reasoning benchmarks, Vision Wormhole improves accuracy by 6.0 percentage points on average over text-mediated MAS and achieves a 1.69× geometric-mean speedup in batch-normalized end-to-end runtime.
Figures & tables
P/R: Gemma-3-4B, C/J: Qwen3-VL-2B
P/R: SmolVLM2-2.2B, C/J: Qwen3-VL-2B
Dataset
Text Acc
Text Time
VW Acc
VW Time
Δ Acc
Speedup
Text Acc
Text Time
VW Acc
VW Time
Δ Acc
Speedup
GSM8K
80.8%
27.3s
77.6%
22.9s
-3.2pp
1.19×
64.3%
63.8s
77.0%
25.6s
+12.7pp
2.49×
ARC-Easy
93.4%
33.1s
90.3%
23.1s
-3.1pp
1.43×
88.6%
51.1s
91.8%
23.6s
+3.2pp
2.17×
ARC-Challenge
86.0%
49.0s
80.5%
30.3s
-5.5pp
1.62×
78.2%
68.5s
81.3%
30.4s
+3.1pp
2.25×
GPQA
29.8%
348.4s
34.9%
172.9s
+5.1pp
2.02×
32.3%
483.3s
33.3%
172.5s
+1.0pp
2.80×
MedQA
53.3%
91.5s
45.0%
82.0s
-8.3pp
1.12×
44.7%
125.0s
49.0%
75.8s
+4.3pp
1.65×
Table 1: Weakly supervised results. We report accuracy (%), batch-normalized runtime (s/query), and differences relative to TextMAS ( Δ Acc in pp; runtime ratio in × ).
Table 2: Main heterogeneous MAS model configurations. Two-backbone configurations alternate backbones across roles.
Dataset
Max. new tokens
Default batch size
GSM8K
2048
12
ARC-Easy
2048
12
ARC-Challenge
2048
12
MedQA
4096
8
MBPP-Plus
4096
8
HumanEval-Plus
4096
8
Appendix
Table 3: Per-dataset generation budgets and default batch sizes. We adopt the same maximum generation budgets as Zou et al. (2025) .
P/R: Gemma-3-4B C/J: Qwen3-VL-2B
P/R: SmolVLM2-2.2B C/J: Qwen3-VL-2B
Dataset
Text
VW
OCR
Text
VW
OCR
GSM8K
80.8% / 27.3s
76.2% / 26.7s
72.8% / 112.9s
64.3% / 63.8s
74.8% / 33.5s
58.0% / 115.0s
ARC-Easy
93.4% / 33.1s
92.4% / 22.0s
82.0% / 113.9s
88.6% / 51.1s
92.3% / 28.2s
73.7% / 113.9s
ARC-Challenge
86.0% / 49.0s
82.1% / 29.5s
72.4% / 113.9s
78.2% / 68.5s
81.7% / 38.2s
64.0% / 112.7s
GPQA
29.8% / 348.4s
39.9% / 174.6s
28.8% / 216.2s
32.3% / 483.3s
37.9% / 225.5s
28.3% / 210.1s
MedQA
53.3% / 91.5s
48.0% / 83.0s
43.3% / 137.4s
44.7% / 125.0s
47.0% / 93.4s
37.0% / 133.6s
Appendix
Table 4: OCR-based image-relay baseline. OCR renders the sender’s generated text as an image, which the receiver reads through its native visual input. Each cell reports accuracy (%) and average wall-clock time (s/query).
Method
Gemma / Qwen
LFM / Qwen
LatentMAS-Hybrid, 16 steps
0.5
23.5
48 steps
0.0
26.0
128 steps
0.0
16.5
192 steps
0.5
18.0
256 steps
0.0
20.0
512 steps
0.0
12.5
Appendix
Table 5: LatentMAS-Hybrid and VW on the same GSM8K subset. Accuracy (%), with correct counts in parentheses. VW uses the main-run generation settings in Appendix B.3 .
Channel
GSM8K
HumanEval+
Vision Wormhole
77.0% (197)
37.2% (61)
CIPHER
65.6% (168)
21.3% (35)
LatentMAS-Hybrid
0 0.0% (0)
0 0.0% (0)
Appendix
Table 6: System accuracy for heterogeneous latent communication channels. Accuracy (%), with correct counts in parentheses. All methods use the Gemma/Qwen role assignment in Table 2 .
Method
GSM8K
MedQA
Qwen + VW
75.8
34.4
Qwen + soft-prefix
60.2
27.0
Δ (pp)
+15.6
+7.4
Appendix
Table 7: Visual interface ablation on Qwen. Accuracy (%); differences in pp.
Intermediate token cap
GSM8K
HumanEval+
256
79.3
34.1
512
80.5
—
1024
78.1
32.9
Task cap
80.5
40.9
Appendix
Table 8: TextMAS with capped intermediate messages. Accuracy (%); “Task cap” uses the per-task limit in Table 3 .
Configuration
TextMAS
Concise TextMAS
VW
SmolVLM2 / Gemma
65.6 (84)
69.5 (89)
82.8 (106)
SmolVLM2 / Qwen
64.8 (83)
50.8 (65)
72.7 (93)
Gemma / Qwen
78.9 (101)
75.0 (96)
74.2 (95)
LFM / Gemma
77.3 (99)
74.2 (95)
85.2 (109)
LFM / Qwen
68.0 (87)
64.8 (83)
74.2 (95)
Four-model pool
57.8 (74)
63.3 (81)
72.7 (93)
Appendix
Table 9: Standard and concise text communication across all six configurations. GSM8K accuracy (%), with correct counts in parentheses. The pooled row aggregates 768 configuration–question evaluations.
Task
n
Text
OCR relay
VW
GSM8K
128
90.6 (116)
92.2 (118)
96.9 (124)
HumanEval+
164
64.0 (105)
44.5 (73)
65.2 (107)
Appendix
Table 10: Communication channels in a parallel MAS. Accuracy (%), with correct counts in parentheses.
Communication
Retrieval accuracy
Text message
34.0±8.7
Vision Wormhole
36.0±4.6
Appendix
Table 11: Key–value retrieval through the Qwen → Gemma interface. Exact-selection accuracy (%), reported as mean ± standard deviation across seeds.
Distractor entries
Text message
Vision Wormhole
0
100.0±0.0
88.9±5.7
1
91.1±4.2
64.4±7.9
4
66.7±7.2
53.3±7.2
16
17.8±9.6
34.4±6.8
64
3.3±4.7
26.7±9.8
Appendix
Table 12: Gemma retrieval as memory size increases. Exact-selection accuracy (%), reported as mean ± standard deviation across seeds.
Direction
Aligned message
Cross-text message
Shuffled-anchor map
Gemma → Qwen
0.227
0.513
0.464
Qwen → Gemma
1.864
6.344
5.408
Appendix
Table 13: Aligned messages reduce reconstruction error. Mean KL divergence ( ↓ ) on held-out texts.
Readout
CKA
Top-1
Top-5
Top-10
Raw pooled representations
0.261
0.20
1.15
2.45
Affine-calibrated
0.840
10.05
17.15
22.15
Chance retrieval
—
0.10
0.50
1.00
Appendix
Table 14: Affine calibration of cross-model representations. CKA and retrieval accuracy (%) on the paired calibration captions.
Decoder input
Gate value
Gaussian message
0.0185±0.0001
Caption-derived message
0.6551±0.0066
Appendix
Table 15: Qwen decoder gate response. Mean gate value and standard deviation across inputs.
Fully correct senders
GSM8K ( n=100 )
HumanEval+ ( n=100 )
AIME 2024 ( n=30 )
Three
39.0
8.0
0.0
Two
45.0
24.0
13.3
One
12.0
53.0
13.3
Zero
4.0
15.0
73.3
Appendix
Table 16: Completeness of sender outputs. Percentage of questions with the indicated number of fully correct intermediate roles.
Latent-space communication allows heterogeneous vision-language model agents to exchange continuous representations without serializing visual and reasoning states into text. Vision Wormhole realizes this approach by translating visual features into a universal latent representation that can be consumed by another model, but every message is transported as a dense tensor of the same size regardless of its content. A fixed-capacity dense tensor therefore need not have a fixed effective information density: some messages may use only a small fraction of the available representational degrees of freedom. This observation suggests that the communication channel may be substantially compressible. We study its redundancy by fitting a post-hoc sparse autoencoder to frozen Vision Wormhole activations and measuring reconstruction, downstream utility, feature reuse, and token-level interventions across nine reasoning benchmarks. Relative to the original float32 transport, a uint16-index/float16-value sparse payload with k=4 active coefficients per token reduces the transmitted bytes by 128x. In a single-run evaluation, the seven-task non-AIME mean accuracy changes from 49.85% to 49.77%. The fitted 4096-element dictionary uses only 50 features, and task-level active sets have a mean pairwise Jaccard similarity of 0.906. These measurements establish strong post-hoc compressibility relative to the original transport, but do not yet isolate the incremental contribution of sparse coding from position selection, reduced precision, low-rank structure, or SAE optimization effects. The results motivate matched-payload comparisons and communication mechanisms whose payload adapts to the information used by each message.
Language-model agents build internal representations of the information they observe and the reasoning they perform. Sharing these representations offers a way to communicate both source information and reasoning across agents. For agents built from different models, this requires aligning their representations while preserving information useful to the receiver. We study this problem through KV-cache communication, examining how an agent uses internal states shared by other agents, with or without direct access to the information that other agents observed. A controlled self-communication study shows that cache pruning causes substantially greater degradation when the receiving agent lacks access to that information. We use this finding to guide dense cross-model cache alignment, combining positional disentanglement and KV-group transformations with reconstruction followed by generation training. Across six directed Qwen3 pairs, aligned caches improve in-domain accuracy over text communication when both agents observe the same input, with fewer estimated inference FLOPs. Experiments with three-agent document sharing and Mistral-to-Qwen transfer further demonstrate that aligned caches can carry information across both multiple separate observations and different model families.
Siyi Chen, Xiaoyan Zhang, Meng Wu +7
University of Michigan · NVIDIA · University of Pennsylvania +2
Multi-agent systems built on large language models have shown strong performance on complex reasoning tasks, yet most work focuses on agent roles and orchestration while treating inter-agent communication as a fixed interface. Latent communication through internal representations such as key-value caches offers a promising alternative to text-based protocols, but existing approaches do not jointly optimize communication with multi-agent reasoning. Therefore we propose DiffMAS, a training framework that treats latent communication as a learnable component of multi-agent systems. DiffMAS performs parameter-efficient supervised training over multi-agent latent trajectories, enabling agents to jointly learn how information should be encoded and interpreted across interactions. Experiments on mathematical reasoning, scientific QA, code generation, and commonsense benchmarks show that DiffMAS consistently improves reasoning accuracy and decoding stability over single-agent inference, text-based multi-agent systems, and prior latent communication methods, achieving 26.7% on AIME24, 20.2% on GPQA-Diamond, and consistent gains across reasoning benchmarks.