Heterogeneous multi-agent systems combine models with different capabilities through a common communication interface. Exchanging internal states directly requires translating between model-specific representations and controlling intermediate computation. We introduce the Vision Wormhole, which repurposes the visual input interface of Vision-Language Models (VLMs) for continuous communication between frozen heterogeneous agents. A Universal Visual Codec encodes each sender's latent rollout into a fixed-size message, maps it through a shared reference space, and decodes received messages into the receiver's image-token span. Per-model codecs and affine reference maps form a hub-and-spoke architecture with O(N) components for N models. Each model learns its codec independently through self-distillation on anchor texts, and shared-anchor alignment enables reuse across communication partners. Across four VLM families, six team configurations, and nine reasoning benchmarks, Vision Wormhole improves accuracy by 6.0 percentage points on average over text-mediated MAS and achieves a 1.69× geometric-mean speedup in batch-normalized end-to-end runtime.
Figures & tables
P/R: Gemma-3-4B, C/J: Qwen3-VL-2B
P/R: SmolVLM2-2.2B, C/J: Qwen3-VL-2B
Dataset
Text Acc
Text Time
VW Acc
VW Time
Δ Acc
Speedup
Text Acc
Text Time
VW Acc
VW Time
Δ Acc
Speedup
GSM8K
80.8%
27.3s
77.6%
22.9s
-3.2pp
1.19×
64.3%
63.8s
77.0%
25.6s
+12.7pp
2.49×
ARC-Easy
93.4%
33.1s
90.3%
23.1s
-3.1pp
1.43×
88.6%
51.1s
91.8%
23.6s
+3.2pp
2.17×
ARC-Challenge
86.0%
49.0s
80.5%
30.3s
-5.5pp
1.62×
78.2%
68.5s
81.3%
30.4s
+3.1pp
2.25×
GPQA
29.8%
348.4s
34.9%
172.9s
+5.1pp
2.02×
32.3%
483.3s
33.3%
172.5s
+1.0pp
2.80×
MedQA
53.3%
91.5s
45.0%
82.0s
-8.3pp
1.12×
44.7%
125.0s
49.0%
75.8s
+4.3pp
1.65×
Table 1: Weakly supervised results. We report accuracy (%), batch-normalized runtime (s/query), and differences relative to TextMAS ( Δ Acc in pp; runtime ratio in × ).
Table 2: Main heterogeneous MAS model configurations. Two-backbone configurations alternate backbones across roles.
Dataset
Max. new tokens
Default batch size
GSM8K
2048
12
ARC-Easy
2048
12
ARC-Challenge
2048
12
MedQA
4096
8
MBPP-Plus
4096
8
HumanEval-Plus
4096
8
Appendix
Table 3: Per-dataset generation budgets and default batch sizes. We adopt the same maximum generation budgets as Zou et al. (2025) .
P/R: Gemma-3-4B C/J: Qwen3-VL-2B
P/R: SmolVLM2-2.2B C/J: Qwen3-VL-2B
Dataset
Text
VW
OCR
Text
VW
OCR
GSM8K
80.8% / 27.3s
76.2% / 26.7s
72.8% / 112.9s
64.3% / 63.8s
74.8% / 33.5s
58.0% / 115.0s
ARC-Easy
93.4% / 33.1s
92.4% / 22.0s
82.0% / 113.9s
88.6% / 51.1s
92.3% / 28.2s
73.7% / 113.9s
ARC-Challenge
86.0% / 49.0s
82.1% / 29.5s
72.4% / 113.9s
78.2% / 68.5s
81.7% / 38.2s
64.0% / 112.7s
GPQA
29.8% / 348.4s
39.9% / 174.6s
28.8% / 216.2s
32.3% / 483.3s
37.9% / 225.5s
28.3% / 210.1s
MedQA
53.3% / 91.5s
48.0% / 83.0s
43.3% / 137.4s
44.7% / 125.0s
47.0% / 93.4s
37.0% / 133.6s
Appendix
Table 4: OCR-based image-relay baseline. OCR renders the sender’s generated text as an image, which the receiver reads through its native visual input. Each cell reports accuracy (%) and average wall-clock time (s/query).
Method
Gemma / Qwen
LFM / Qwen
LatentMAS-Hybrid, 16 steps
0.5
23.5
48 steps
0.0
26.0
128 steps
0.0
16.5
192 steps
0.5
18.0
256 steps
0.0
20.0
512 steps
0.0
12.5
Appendix
Table 5: LatentMAS-Hybrid and VW on the same GSM8K subset. Accuracy (%), with correct counts in parentheses. VW uses the main-run generation settings in Appendix B.3 .
Channel
GSM8K
HumanEval+
Vision Wormhole
77.0% (197)
37.2% (61)
CIPHER
65.6% (168)
21.3% (35)
LatentMAS-Hybrid
0 0.0% (0)
0 0.0% (0)
Appendix
Table 6: System accuracy for heterogeneous latent communication channels. Accuracy (%), with correct counts in parentheses. All methods use the Gemma/Qwen role assignment in Table 2 .
Method
GSM8K
MedQA
Qwen + VW
75.8
34.4
Qwen + soft-prefix
60.2
27.0
Δ (pp)
+15.6
+7.4
Appendix
Table 7: Visual interface ablation on Qwen. Accuracy (%); differences in pp.
Intermediate token cap
GSM8K
HumanEval+
256
79.3
34.1
512
80.5
—
1024
78.1
32.9
Task cap
80.5
40.9
Appendix
Table 8: TextMAS with capped intermediate messages. Accuracy (%); “Task cap” uses the per-task limit in Table 3 .
Configuration
TextMAS
Concise TextMAS
VW
SmolVLM2 / Gemma
65.6 (84)
69.5 (89)
82.8 (106)
SmolVLM2 / Qwen
64.8 (83)
50.8 (65)
72.7 (93)
Gemma / Qwen
78.9 (101)
75.0 (96)
74.2 (95)
LFM / Gemma
77.3 (99)
74.2 (95)
85.2 (109)
LFM / Qwen
68.0 (87)
64.8 (83)
74.2 (95)
Four-model pool
57.8 (74)
63.3 (81)
72.7 (93)
Appendix
Table 9: Standard and concise text communication across all six configurations. GSM8K accuracy (%), with correct counts in parentheses. The pooled row aggregates 768 configuration–question evaluations.
Task
n
Text
OCR relay
VW
GSM8K
128
90.6 (116)
92.2 (118)
96.9 (124)
HumanEval+
164
64.0 (105)
44.5 (73)
65.2 (107)
Appendix
Table 10: Communication channels in a parallel MAS. Accuracy (%), with correct counts in parentheses.
Communication
Retrieval accuracy
Text message
34.0±8.7
Vision Wormhole
36.0±4.6
Appendix
Table 11: Key–value retrieval through the Qwen → Gemma interface. Exact-selection accuracy (%), reported as mean ± standard deviation across seeds.
Distractor entries
Text message
Vision Wormhole
0
100.0±0.0
88.9±5.7
1
91.1±4.2
64.4±7.9
4
66.7±7.2
53.3±7.2
16
17.8±9.6
34.4±6.8
64
3.3±4.7
26.7±9.8
Appendix
Table 12: Gemma retrieval as memory size increases. Exact-selection accuracy (%), reported as mean ± standard deviation across seeds.
Direction
Aligned message
Cross-text message
Shuffled-anchor map
Gemma → Qwen
0.227
0.513
0.464
Qwen → Gemma
1.864
6.344
5.408
Appendix
Table 13: Aligned messages reduce reconstruction error. Mean KL divergence ( ↓ ) on held-out texts.
Readout
CKA
Top-1
Top-5
Top-10
Raw pooled representations
0.261
0.20
1.15
2.45
Affine-calibrated
0.840
10.05
17.15
22.15
Chance retrieval
—
0.10
0.50
1.00
Appendix
Table 14: Affine calibration of cross-model representations. CKA and retrieval accuracy (%) on the paired calibration captions.
Decoder input
Gate value
Gaussian message
0.0185±0.0001
Caption-derived message
0.6551±0.0066
Appendix
Table 15: Qwen decoder gate response. Mean gate value and standard deviation across inputs.
Fully correct senders
GSM8K ( n=100 )
HumanEval+ ( n=100 )
AIME 2024 ( n=30 )
Three
39.0
8.0
0.0
Two
45.0
24.0
13.3
One
12.0
53.0
13.3
Zero
4.0
15.0
73.3
Appendix
Table 16: Completeness of sender outputs. Percentage of questions with the indicated number of fully correct intermediate roles.