cs.CVJun 23, 2026

Latent Visual States for Efficient Multimodal Reasoning

Authors: Xiuwei ChenWentao HuYongxin WangZisheng ChenLikui ZhangKun XiangJianhua HanHui-Ling Zhen+3 more

Organizations: Sun Yat-sen University · The Hong Kong Polytechnic University · MBZUAI · Yinwang Intelligent Technology Co. Ltd. · Huawei Noah’s Ark Lab

Abstract

The integration of visual evidence has significantly enhanced the capabilities of large multimodal models. However, this integration predominantly relies on generating discrete outputs (etc., code or box coordinates) to invoke external tools, a process that introduces rigid dependencies and substantial latency. To overcome these limitations, we propose {EVA} (LatEnt Visual StAtes), a novel framework that natively generates continuous latent visual representations. These internal representations manifest as an adaptive sequence of Latent_slot tokens, serving as intermediate visual thoughts during the reasoning process. These Latent_slot tokens are then trained end-to-end with the discrete text tokens. This co-optimization, notably, causes extreme policy deviation in the 'transition window' following the Latent_slot tokens. We develop D-GSPO (Decouple-GSPO) to target this root cause by decoupling the optimization of latent and discrete components. To support SFT, we construct EVA-230K, a high-quality text-image interleaved CoT dataset encompassing a diverse range of real-world scenes, documents, charts and OCR tasks. Extensive experiments across multiple benchmarks confirm that EVA achieves significant performance gains while enhancing inference efficiency.

Explore similar work

CardsList