cs.MAOct 6, 2026

Visual Orchestration Tax in Agentic VLM Pipelines: Auditing and Certifying Visual Evidence Reuse

Authors: Lingteng Zeng

Organizations: Faculty of Engineering, The Chinese University of Hong Kong Hong Kong, China

Abstract

Agentic VLM pipelines increasingly pass the same static visual evidence through multiple specialist agents and tools. This design creates an orchestration-level redundancy mode: semantically unchanged images are repeatedly reconstructed as image-conditioned requests at the VLM API boundary. We call this phenomenon visual orchestration tax and develop a measurement-to-certification framework for visual evidence reuse in agentic VLM pipelines. The audit side defines M1trace\mathrm{M1}_{\mathrm{trace}} to count raw visual-evidence touches and M2 to measure structural touch redundancy, with query-level distributions, bootstrap confidence intervals, and paired quality tests. Across SeeingEye and MAMMQA on chart, document, general-VQA, and multi-modal-QA tasks, audits reveal 66.8-75.6% visual-evidence touch redundancy, and every audited query exceeds the predefined gate. The certification side introduces SharedVisCache, a contract-aware evidence reuse hook keyed by image content, preprocessing fingerprint, and encoder assumptions. On SeeingEye, contract validation certifies 75.0-75.5% repeated touches as reusable while preserving 350/350 output strings and ΔM5=0Δ\mathrm{M5}{=}0. At the physical layer, certified hits reduce FvisionF_{\mathrm{vision}} from 800 to 200 in ChartQA-200 trace replay and from 200 to 50 inside live SeeingEye translator-stage physical integration, preserving 800/800 replay strings and 200/200 integrated call outputs. The results position visual reuse as a measurable, behavior-preserving property of agent orchestration and define an agent-layer contract that makes backend prefix or token reuse semantically interpretable.

Figures & tables

Explore similar work

CardsList
  1. LAVE: Latent Visual Evidence-Enhanced Planning for Video Tool-use Agents

    Aug 5, 2026Zijian Wang, Junnan Zhu, Rongzhen Li +7Video AgentVideo Understanding

  2. Seeing Before Agreeing: Aligning Multi-Agent Consensus with Visual Evidence

    May 29, 2026Yuhan Wang, Shuochen Chang, Yalin Feng +8Knowledge-Based Visual Question AnsweringMultimodal Agents

  3. Are VLMs Seeing or Just Saying? Uncovering the Illusion of Visual Re-examination

    May 15, 2026Chufan Shi, Cheng Yang, Yaokang Wu +4Vision-Language Foundation ModelsIllusions