cs.HCSep 30, 2026

Faithful Chart Generation for Multimodal Deep Research: Frame-Evidence Co-Adaptation

Authors: Yuxin Yue, Yingchen Zhang, Ruqing Zhang, Jiafeng Guo, Maarten de Rijke, Xueqi Cheng

Organizations: State Key Laboratory of AI Safety Institute of Computing Technology, Chinese Academy of Sciences University of Chinese Academy of Sciences Beijing, China · University of Amsterdam Amsterdam, The Netherlands

Abstract

Analytical charts in multimodal deep research encode quantitative claims, requiring every visualized value to be faithfully grounded in supporting evidence. Unlike retrieved images that mainly provide contextual information, charts require numerical fidelity: visualized values should not only match retrieved evidence quantitatively but also preserve its original meaning and scope. However, achieving such fidelity remains challenging because current systems usually construct visualization plans before knowing what quantitative evidence can actually be retrieved from the web. As a result, predefined plans may require entities, temporal ranges, or comparison dimensions that the retrieved evidence only partially supports. Existing approaches mainly address this issue through post-hoc verification after chart plans are fixed, enabling unsupported values to be identified but leaving the underlying visual frames unchanged. To address this challenge, we propose Frame-Evidence Co-Adaptation (FECA), an evidence-adaptive visual planning framework for multimodal deep research. Inspired by the bidirectional sensemaking process in Data-Frame Theory, FECA models chart generation as an iterative interaction between visual frames and retrieved evidence. Each visual frame is adaptive: the frame guides evidence acquisition, while retrieved evidence determines whether the frame should be accepted, revised, or dropped before rendering. By coupling visualization planning with evidence availability, FECA shifts chart generation from fixed-plan verification to adaptive evidence-grounded visual reasoning. Experiments on 100 real-world research topics show that FECA substantially improves numerical fidelity while preserving report quality and chart utility.

Figures & tables

Explore similar work

May 13, 2026cs.CV

ViDR: Grounding Multimodal Deep Research Reports in Source Visual Evidence

Recent deep research systems have improved the ability of large language models to produce long, grounded reports through iterative retrieval and reasoning. However, most text-centered systems rely mainly on textual evidence, while multimodal systems often retrieve images only weakly or generate charts themselves, leaving source figures underused as evidence. We present ViDR, a multimodal deep research framework that grounds long-form reports in source figures. ViDR treats source figures as retrievable, interpretable, routable, and verifiable evidence objects, while still generating analytical charts when needed. It builds an evidence-indexed outline linking claims to textual and visual evidence, refines noisy web images into source-figure evidence atoms through context-aware filtering, outline-aware reranking, and VLM-based visual analysis, and generates each section with section-specific evidence. ViDR further validates visual references to reduce hallucinated or misplaced figures. We also introduce MMR Bench+, a benchmark for evaluating visual evidence use in deep research reports, covering source-figure retrieval, placement, interpretation, verifiability, and analytical chart generation. Experiments show that ViDR improves overall report quality, source-figure integration, and verifiability over strong commercial and open-source baselines. These results suggest that source visual evidence is important for multimodal deep research, as it strengthens evidential grounding, visual support, and report verifiability.
May 3, 2026cs.CV

Chart-FR1: Visual Focus-Driven Fine-Grained Reasoning on Dense Charts

Multimodal large language models (MLLMs) have shown considerable potential in chart understanding and reasoning tasks. However, they still struggle with high information density (HID) charts characterized by multiple subplots, legends, and dense annotations due to three major challenges: (1) limited fine-grained perception results in the omission of critical visual cues; (2) redundant or noisy visual information undermines the performance of multimodal reasoning; (3) lack of adaptive deep reasoning relative to the amount of visual information. To tackle these challenges, we present a novel focus-driven fine-grained chart reasoning model, Chart-FR1, to improve perception, focusing efficiency, and adaptive deep reasoning on HID charts. Specifically, we propose Focus-CoT, a visual focusing chain-of-thought that enhances fine-grained perception by explicitly linking reasoning steps to key visual cues, such as local image regions and OCR signals. Building on this, we introduce Focus-GRPO, a focus-driven reinforcement learning algorithm with an information-efficiency reward that compresses redundant visual information for efficient focusing, and an adaptive KL penalty mechanism that enables flexible control over reasoning depth as more visual cues are discovered. Furthermore, to fill the gap in benchmarks for HID charts, we build HID-Chart, a challenging benchmark with an information-density metric designed to evaluate fine-grained chart reasoning capabilities. Extensive experiments on multiple chart benchmarks demonstrate that Chart-FR1 outperforms state-of-the-art MLLMs in chart understanding and reasoning. Code is available at https://github.com/phkhub/Chart-FR1.
May 5, 2026cs.CV

Hierarchical Visual Agent: Managing Contexts in Joint Image-Text Space for Advanced Chart Reasoning

Advanced chart question answering requires both precise perception of small visual elements and multi-step reasoning across several subplots. While existing MLLMs are strong at understanding single plots, they often struggle with multi-step reasoning across multiple subplots. We propose HierVA, a hierarchical visual agent framework for chart reasoning that iteratively constructs and updates a working context in a joint image--text space. A high-level manager generates plans and maintains a compact context containing only key information, while specialized workers perform reasoning, gather evidence, and return results. In particular, the agent maintains separate visual and textual contexts, using a zoom-in tool to restrict the visual context. Experiments on the CharXiv reasoning subset demonstrate consistent improvements over strong multimodal baselines, and ablation studies verify that hierarchical architecture, scoped visual context, and distilled context contribute complementary gains.