cs.CLFeb 17, 2026

Vision Wormhole: Latent-Space Communication in Heterogeneous Multi-Agent Systems

Authors: Xiaoze Liu, Ruowang Zhang, Weichen Yu, Siheng Xiong, Liu He, Feijie Wu, Hoin Jung, Matt Fredrikson, +2 more

Organizations: Purdue University · Contextual AI · Carnegie Mellon University · Georgia Institute of Technology

Abstract

Heterogeneous multi-agent systems combine models with different capabilities through a common communication interface. Exchanging internal states directly requires translating between model-specific representations and controlling intermediate computation. We introduce the Vision Wormhole, which repurposes the visual input interface of Vision-Language Models (VLMs) for continuous communication between frozen heterogeneous agents. A Universal Visual Codec encodes each sender's latent rollout into a fixed-size message, maps it through a shared reference space, and decodes received messages into the receiver's image-token span. Per-model codecs and affine reference maps form a hub-and-spoke architecture with O(N)O(N) components for NN models. Each model learns its codec independently through self-distillation on anchor texts, and shared-anchor alignment enables reuse across communication partners. Across four VLM families, six team configurations, and nine reasoning benchmarks, Vision Wormhole improves accuracy by 6.0 percentage points on average over text-mediated MAS and achieves a 1.69×\times geometric-mean speedup in batch-normalized end-to-end runtime.

Figures & tables

Appendix figures & tables15 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

Aug 10, 2026cs.AI

Post-Hoc Sparse Coding of Latent Communication Between Vision-Language Model Agents

Latent-space communication allows heterogeneous vision-language model agents to exchange continuous representations without serializing visual and reasoning states into text. Vision Wormhole realizes this approach by translating visual features into a universal latent representation that can be consumed by another model, but every message is transported as a dense tensor of the same size regardless of its content. A fixed-capacity dense tensor therefore need not have a fixed effective information density: some messages may use only a small fraction of the available representational degrees of freedom. This observation suggests that the communication channel may be substantially compressible. We study its redundancy by fitting a post-hoc sparse autoencoder to frozen Vision Wormhole activations and measuring reconstruction, downstream utility, feature reuse, and token-level interventions across nine reasoning benchmarks. Relative to the original float32 transport, a uint16-index/float16-value sparse payload with k=4 active coefficients per token reduces the transmitted bytes by 128x. In a single-run evaluation, the seven-task non-AIME mean accuracy changes from 49.85% to 49.77%. The fitted 4096-element dictionary uses only 50 features, and task-level active sets have a mean pairwise Jaccard similarity of 0.906. These measurements establish strong post-hoc compressibility relative to the original transport, but do not yet isolate the incremental contribution of sparse coding from position selection, reduced precision, low-rank structure, or SAE optimization effects. The results motivate matched-payload comparisons and communication mechanisms whose payload adapts to the information used by each message.
Jun 11, 2026cs.MA

See What I See, Know What I Think: Dense Latent Communication Across Heterogeneous Agents

Language-model agents build internal representations of the information they observe and the reasoning they perform. Sharing these representations offers a way to communicate both source information and reasoning across agents. For agents built from different models, this requires aligning their representations while preserving information useful to the receiver. We study this problem through KV-cache communication, examining how an agent uses internal states shared by other agents, with or without direct access to the information that other agents observed. A controlled self-communication study shows that cache pruning causes substantially greater degradation when the receiving agent lacks access to that information. We use this finding to guide dense cross-model cache alignment, combining positional disentanglement and KV-group transformations with reconstruction followed by generation training. Across six directed Qwen3 pairs, aligned caches improve in-domain accuracy over text communication when both agents observe the same input, with fewer estimated inference FLOPs. Experiments with three-agent document sharing and Mistral-to-Qwen transfer further demonstrate that aligned caches can carry information across both multiple separate observations and different model families.
Apr 23, 2026cs.AI

Learning to Communicate: Toward End-to-End Optimization of Multi-Agent Language Systems

Multi-agent systems built on large language models have shown strong performance on complex reasoning tasks, yet most work focuses on agent roles and orchestration while treating inter-agent communication as a fixed interface. Latent communication through internal representations such as key-value caches offers a promising alternative to text-based protocols, but existing approaches do not jointly optimize communication with multi-agent reasoning. Therefore we propose DiffMAS, a training framework that treats latent communication as a learnable component of multi-agent systems. DiffMAS performs parameter-efficient supervised training over multi-agent latent trajectories, enabling agents to jointly learn how information should be encoded and interpreted across interactions. Experiments on mathematical reasoning, scientific QA, code generation, and commonsense benchmarks show that DiffMAS consistently improves reasoning accuracy and decoding stability over single-agent inference, text-based multi-agent systems, and prior latent communication methods, achieving 26.7% on AIME24, 20.2% on GPQA-Diamond, and consistent gains across reasoning benchmarks.