cs.HCSep 30, 2026

Faithful Chart Generation for Multimodal Deep Research: Frame-Evidence Co-Adaptation

Authors: Yuxin Yue, Yingchen Zhang, Ruqing Zhang, Jiafeng Guo, Maarten de Rijke, Xueqi Cheng

Organizations: State Key Laboratory of AI Safety Institute of Computing Technology, Chinese Academy of Sciences University of Chinese Academy of Sciences Beijing, China · University of Amsterdam Amsterdam, The Netherlands

Abstract

Analytical charts in multimodal deep research encode quantitative claims, requiring every visualized value to be faithfully grounded in supporting evidence. Unlike retrieved images that mainly provide contextual information, charts require numerical fidelity: visualized values should not only match retrieved evidence quantitatively but also preserve its original meaning and scope. However, achieving such fidelity remains challenging because current systems usually construct visualization plans before knowing what quantitative evidence can actually be retrieved from the web. As a result, predefined plans may require entities, temporal ranges, or comparison dimensions that the retrieved evidence only partially supports. Existing approaches mainly address this issue through post-hoc verification after chart plans are fixed, enabling unsupported values to be identified but leaving the underlying visual frames unchanged. To address this challenge, we propose Frame-Evidence Co-Adaptation (FECA), an evidence-adaptive visual planning framework for multimodal deep research. Inspired by the bidirectional sensemaking process in Data-Frame Theory, FECA models chart generation as an iterative interaction between visual frames and retrieved evidence. Each visual frame is adaptive: the frame guides evidence acquisition, while retrieved evidence determines whether the frame should be accepted, revised, or dropped before rendering. By coupling visualization planning with evidence availability, FECA shifts chart generation from fixed-plan verification to adaptive evidence-grounded visual reasoning. Experiments on 100 real-world research topics show that FECA substantially improves numerical fidelity while preserving report quality and chart utility.

Figures & tables

Explore similar work

CardsList
  1. ViDR: Grounding Multimodal Deep Research Reports in Source Visual Evidence

    May 13, 2026Zhuofan Shi, Peilun Jia, Baoqin Sun +4Visual EvidenceDeep Research

  2. Chart-FR1: Visual Focus-Driven Fine-Grained Reasoning on Dense Charts

    May 3, 2026Hongkun Pan, Yuwei Wu, Wanyi Hong +8ChartMultimodal Reasoning

  3. Hierarchical Visual Agent: Managing Contexts in Joint Image-Text Space for Advanced Chart Reasoning

    May 5, 2026Qihua Dong, Ruozhen He, Junwen Chen +4ChartVisual Reasoning