cs.AISep 28, 2026

FigAct: Turning Scientific Figures into Active Canvases for Explanation

Authors: Shishi Xiao, Zichao Wang, Alexa Siu, David H. Laidlaw, Jennifer Healey

Organizations: Brown University · Adobe Research

Abstract

Scientific figures are designed to communicate information visually, yet MLLMs typically explain them by translating their visual content back into text. This requires readers to manually map the resulting explanations back to the figure. Inspired by how people present visual information, we introduce FigAct, a framework that transforms static scientific figures into question-conditioned visual presentations by acting directly on their existing graphical elements. Like a human presenter, FigAct generates a sequence of short narrations, grounds each narration in the corresponding visual evidence, and applies visual actions to guide the viewer's attention. We develop a hierarchical search strategy for efficient element localization, reducing token usage by approximately 40×\times. We further train FigAct-8B using three task-specific rewards for grounding accuracy, search efficiency, and rendering quality. We further build a human-verified benchmark from figures in real-world scientific papers to evaluate the ability of MLLMs to generate grounded visual explanations. Our results demonstrate the effectiveness of FigAct and show that treating scientific figures as presentation canvases makes explanations clearer and easier to follow.

Figures & tables

Appendix figures & tables16 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Helping Figures Tell their Story! Paper-Grounded Video Generation Explaining Complex Scientific Figures

    Jun 10, 2026Ishani Mondal, Javad Baghirov, Jordan Boyd-GraberLong Video GenerationScientific Figure

  2. Diagram-MMU: A Multi-Modal Benchmark for Scientific Diagrams

    Aug 12, 2026Weihao Bo, Shan Zhang, Yanpeng Sun +7DiagramsDiagnostic Benchmark