cs.AIOct 5, 2026

CVIF: A Criticality-Driven Visual Intervention Framework for Geometric Diagram Understanding in MLLMs

Authors: Jiahui Kang, Bifan Wei, Lingling Zhang, Tianwen Jiang, Qiuyong Xiao, Jihong Zhang, Jun Liu

Organizations: School of Computer Science and Technology, Xi’an Jiaotong University · Tencent Hy AI Data

Abstract

Despite significant progress in visual tasks by Multimodal Large Language Models (MLLMs), geometric diagram understanding remains challenging due to the presence of sparse visual cues and ambiguous symbol-primitive associations. MLLMs may therefore rely on textual priors, producing interpretations that conflict with visual evidence. We introduce the training-free Criticality-Driven Visual Intervention Framework (CVIF), an inference-time method that localizes critical layers and executes visual interventions during the transition from evidence aggregation to semantic decoding. At these layers, a Geometry-Constrained Local Relation Reconstruction (GCLR) module selects and weights vertex-centered visual evidence, while an Adaptive Visual Steering Operator (AVSO) redistributes attention mass toward the selected tokens. Experiments on PGPS9K and PGDP5K show that CVIF raises Overall F1 from 77.85 to 85.58 and from 75.23 to 82.84, respectively, establishing a novel inference-time visual intervention paradigm.

Figures & tables

Explore similar work

CardsList
  1. PGT: Procedurally Generated Tasks for improving visual grounding in MLLMs

    May 22, 2026Rim Assouel, Amir Bar, Michal Drozdzal +1Weak Visual GroundingFine-Grained Perception

  2. Seeing or Knowing? Visual Context Sensitivity in Multimodal Large Language Models

    Jul 28, 2026Jiaang Li, Chengzu Li, Zhaochong An +4Multimodal Large Language ModelsRecent Vision-Language Models

  3. Vision Inference Former: Sustaining Visual Consistency in Multimodal Large Language Models

    May 18, 2026Xinpeng Dong, Min Zhang, Kairong Han +3Recent Vision-Language ModelsMultimodal Large Language Models