cs.CLOct 1, 2026

A Citation-Grounded Benchmark for Trustworthy Earnings Call Transcript Analysis with Large Language Models

Authors: Yingzhu Zhao, Vlad Pandelea, Han Yuan, Bo Hu, Wuqiong Luo, Li Zhang, Zheng Ma

Organizations: Global Decision Science, American Express · Singapore Decision Science Center of Excellence, American Express, 1 Marina Boulevard, 018989, Singapore.

Abstract

Large language models (LLMs) have been increasingly used for financial document analysis, including earnings call transcripts (ECTs). Beyond generating standalone claims, users increasingly prefer grounded analyses that pair claims with verifiable citations from source documents to enable independent validation. However, evaluating such analytical claims typically requires extensive expert annotation, which is costly and difficult to scale, and real-world financial analysis commonly involves long context-question-answer triplets, further increasing task complexity. To address these challenges and benchmark the current landscape of grounded analysis by LLMs, we propose a numeric evidence evaluation method that enables groundedness assessment without reliance on expert annotation. We also introduce an automated dataset construction pipeline and construct ECTs-100 from the top 100 constituents of the S&P 500 to support benchmark of both groundedness and correctness. In addition, we examine conscious incompetence, a practical failure mode in financial analysis in which LLMs must detect when available evidence is insufficient and refrain from producing unsupported hallucinations. Empirical results show that LLMs perform well in groundedness but face notable limitations in correctness, with informational insufficiency presenting an additional challenge.

Figures & tables

Explore similar work

CardsList
  1. Are the Financial Reasoning from LLMs Credible? A Real World Test over Long-Horizon Statements

    Jul 22, 2026Xinke Tong, Xuanming Zhang, Tianyi Tang +10Long-Horizon Task PlanningFidelity

  2. No Free Labels: Limitations of LLM-as-a-Judge Without Human Grounding

    Mar 7, 2025Michael Krumdick, Charles Lovering, Varshini Reddy +2Llm-As-A-JudgeGrading