cs.LGSep 29, 2026

MM-FinEval: A Multi-Task Multimodal Benchmark for Real-World Financial Forecasting

Authors: Dong Shu, Yanguang Liu, Huopu Zhang, Saisai Hu, Haiyan Zhao, Hekun Huang, Mengnan Du

Organizations: Northwestern University · New Jersey Institute of Technology · Georgia Institute of Technology · Pace University · The Chinese University of Hong Kong, Shenzhen

Abstract

Financial forecasting from earnings conference calls requires models to reason over complex corporate disclosures, market expectations, and subtle communication signals. However, existing financial benchmarks are often limited to unimodal inputs or single-task settings, making it difficult to evaluate whether multimodal large language models (LLMs) can support real-world financial analysis. In this paper, we introduce MM-FinEval, a novel benchmark designed to evaluate multimodal LLMs across multiple financial tasks. MM-FinEval spans a diverse timeline from 2019 to 2022. The entire proposed dataset contains 2,045 S&P 500 conference earning calls as inputs and 12 financial task labels as outputs. Each input contains three modalities: a word-to-word text transcript of the earning call, the corresponding presentation slides used during the call, and the entire audio recording. To establish a rigorous evaluation framework, we analyze 19 baseline models across three distinct model categories: Image-Text, Audio-Text, and Any-to-Any configurations. We observe that small-size Any-to-Any models processing all three modalities achieve strong performance, even when compared against larger proprietary models restricted to two-modality inputs. This indicates that our tri-modal dataset design introduces useful, non-redundant information. These results validate that text, audio, and visual data serve as important, complementary signals that mimic the decision-making process of expert human analysts.

Figures & tables

Appendix figures & tables2 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Benchmarking Large Vision-Language Models on CFMME: A Comprehensive Chinese Financial Multimodal Evaluation Dataset

    May 28, 2026Qian Chen, Xianyin Zhang, Yanzhi Liu +3Cross-Modal

  2. A Citation-Grounded Benchmark for Trustworthy Earnings Call Transcript Analysis with Large Language Models

    Oct 1, 2026Yingzhu Zhao, Vlad Pandelea, Han Yuan +4Financial Question AnsweringHuman-Annotated Benchmark