cs.CLFeb 17, 2026

Vision Wormhole: Latent-Space Communication in Heterogeneous Multi-Agent Systems

Authors: Xiaoze Liu, Ruowang Zhang, Weichen Yu, Siheng Xiong, Liu He, Feijie Wu, Hoin Jung, Matt Fredrikson, +2 more

Organizations: Purdue University · Contextual AI · Carnegie Mellon University · Georgia Institute of Technology

Abstract

Heterogeneous multi-agent systems combine models with different capabilities through a common communication interface. Exchanging internal states directly requires translating between model-specific representations and controlling intermediate computation. We introduce the Vision Wormhole, which repurposes the visual input interface of Vision-Language Models (VLMs) for continuous communication between frozen heterogeneous agents. A Universal Visual Codec encodes each sender's latent rollout into a fixed-size message, maps it through a shared reference space, and decodes received messages into the receiver's image-token span. Per-model codecs and affine reference maps form a hub-and-spoke architecture with O(N)O(N) components for NN models. Each model learns its codec independently through self-distillation on anchor texts, and shared-anchor alignment enables reuse across communication partners. Across four VLM families, six team configurations, and nine reasoning benchmarks, Vision Wormhole improves accuracy by 6.0 percentage points on average over text-mediated MAS and achieves a 1.69×\times geometric-mean speedup in batch-normalized end-to-end runtime.

Figures & tables

Appendix figures & tables15 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. See What I See, Know What I Think: Dense Latent Communication Across Heterogeneous Agents

    Jun 11, 2026Siyi Chen, Xiaoyan Zhang, Meng Wu +7Kv-Cache ManagementLatent Representation Alignment