cs.AIOct 7, 2026

The Attribution Blind Spot: Layerwise Trajectory Diagnostics for Source Reliance in Retrieval-Augmented Language Models

Authors: Zhe Yu, Wenpeng Xing, Yunzhao Wei, Bo Yang, Chen Ye, Gaolei Li, Meng Han

Organizations: Zhejiang University · Binjiang Institute of Zhejiang University · East China Normal University · National FinTech Evaluation Center (Bank Card Testing Center) · Hangzhou Dianzi University · Shanghai Jiaotong University

Abstract

A retrieval-augmented model can match a document without relying on it. Controlled knowledge conflicts make source choice observable and let us ask a second question that prediction alone cannot answer: which internal-state properties define useful intervention directions? We study paired hidden-state changes with Latent Trajectory Shift (LTS), a signed projection onto a training-fitted first principal component (PC1), and keep verified training exposure separate from behavioral source choice. Across the evaluated conflicts, state-change magnitude is often the stronger predictor, whereas signed PC1 is the stronger selective controller: equal-norm interventions change source preference while better preserving non-target behavior, and the frozen direction transfers across the tested datasets and aligned model pairs. A same-system OLMo study further combines positive choice and control results with inconclusive exposure detection at the achieved power. The central result is a separation: representations that diagnose what a model will choose need not be the representations that best control that choice.

Figures & tables

Appendix figures & tables21 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. The Attribution Blind Spot: Detecting When Language Models Rely on Memory Rather Than Retrieved Context

    May 26, 2026Zhe Yu, Wenpeng Xing, Yunzhao Wei +4Blind SpotsCognitive Science

  2. Belief-Trajectory Energy: Measuring the Path to a Prediction

    Oct 4, 2026Jiahao Ying, Wei Tang, Boxian Ai +7

  3. The Fellowship of the Query: Learning Retrieval Actions

    Sep 23, 2026Mohammed Al-Maamari, Saber Zerhoudi, Michael Granitzer +1Fine-Grained ActionsModel Fine-Tuning