cs.CLAug 25, 2026

Reading Is Not Using: Retrieval, Judgment, and the Design of AI Financial Research Workflows

Authors: Miao Liu, Zhizhe Liu

Abstract

Large language models (LLMs) are increasingly deployed as AI analysts to process financial disclosures and support AI-assisted investment decisions. Yet such systems are usually evaluated by what they can retrieve, not whether retrieved information affects their judgments. We identify a retrieval-integration gap in long-context financial analysis. Holding focal-firm information fixed and varying only unrelated context from 2,000 to 128,000 tokens, we find that a risk disclosure's influence on investment judgments falls to the experimental noise floor even as direct retrieval remains accurate. The pattern replicates across model families and judgment tasks and in experiments removing real disclosures from actual 10-K filings. More capable models postpone but do not eliminate the gap. Causal memory interventions show that compressed summaries and source-text lookup jointly transmit disclosures into judgments. Workflow architecture determines whether this transmission succeeds: chunk-and-summarize pipelines evict relevant information, whereas a targeted, structured restatement adjacent to the decision restores its influence. AI analyst performance is therefore jointly determined by model capability and workflow architecture. Retrieval-based evaluations can certify systems whose investment judgments ignore information they demonstrably retrieved.

Explore similar work

CardsList
  1. The Analyst in the Prompt: Role, Retrieval, and Memory Biases in LLM Financial Analysis

    Sep 2, 2026Ahmed Asaad, Amr Mohamed, Yang Zhang +1LLM EvaluationLLM Auditing

  2. When Summaries Distort Decisions: Information Fidelity in LLM-Compressed Financial Analysis

    Jun 28, 2026Hoyoung Lee, Suhwan Park, Seunghan Lee +15LLM CompressionLLM Auditing

  3. Fin-Bias: Comprehensive Evaluation for LLM Decision-Making under human bias in Finance Domain

    May 9, 2026Xiaoyu Hu, Jinman ZhaoLLM EvaluationSocial Bias in Language Models