The Attribution Blind Spot: Layerwise Trajectory Diagnostics for Source Reliance in Retrieval-Augmented Language Models
Organizations: Zhejiang University · Binjiang Institute of Zhejiang University · East China Normal University · National FinTech Evaluation Center (Bank Card Testing Center) · Hangzhou Dianzi University · Shanghai Jiaotong University
Abstract
A retrieval-augmented model can match a document without relying on it. Controlled knowledge conflicts make source choice observable and let us ask a second question that prediction alone cannot answer: which internal-state properties define useful intervention directions? We study paired hidden-state changes with Latent Trajectory Shift (LTS), a signed projection onto a training-fitted first principal component (PC1), and keep verified training exposure separate from behavioral source choice. Across the evaluated conflicts, state-change magnitude is often the stronger predictor, whereas signed PC1 is the stronger selective controller: equal-norm interventions change source preference while better preserving non-target behavior, and the frozen direction transfers across the tested datasets and aligned model pairs. A same-system OLMo study further combines positive choice and control results with inconclusive exposure detection at the achieved power. The central result is a separation: representations that diagnose what a model will choose need not be the representations that best control that choice.
Figures & tables
| Arm | What it tests | Rep. |
|---|---|---|
| Train-only PC1 | Signed source direction | 1 |
| Candidate gradient | Direct answer steering | 1 |
| L2 radial | Magnitude without PC1 | 1 |
| Shuffled-layer PC1 | Layer specificity | 1 |
| Isotropic random | Direction specificity | 20 |
| PC1-orthogonal random | PC1 specificity | 20 |
| Test | Endpoint | Estimate [95% CI] | Result |
|---|---|---|---|
| Exposure | ROC–AUC | 0.508 [0.482, 0.534] | Inconclusive |
| Choice | PR–AUC | 0.718 [0.667, 0.769] | Positive |
| Control | 0.392 [0.344, 0.440] | Positive | |
| Transfer | 0.384 [0.336, 0.432] | Positive |
| Context condition | Parametric (%) | Contextual (%) | Other (%) |
|---|---|---|---|
| Congruent | 99.1 | 0.0 | 0.9 |
| Conflict | 31.0 | 63.7 | 5.3 |
| Unrelated, matched | 88.5 | 7.1 | 4.4 |
| Shuffled conflict | 60.2 | 31.9 | 8.0 |
| Paraphrased conflict | 34.5 | 64.6 | 0.9 |
| Mixed | 64.6 | 7.1 | 28.3 |
| Model / items | Method | ROC–AUC | AP | AP 95% CI |
|---|---|---|---|---|
| Llama | LTS | 0.859 | 0.311 | [0.206, 0.444] |
| Supervised LTS | 0.880 | 0.415 | [0.301, 0.546] | |
| L2 magnitude | 0.978 | 0.741 | [0.615, 0.862] | |
| Answer margin | 0.999 | 0.989 | [0.973, 0.998] | |
| Qwen | LTS | 0.903 | 0.716 | [0.586, 0.828] |
| Supervised LTS | 0.912 | 0.631 | [0.502, 0.748] |
| Model / items | Block | Toward parametric | Toward contextual | Controls | ||
|---|---|---|---|---|---|---|
| Effect | 95% CI | Effect | 95% CI | passed | ||
| Llama | Early | 0.018 | [-0.063, 0.100] | 0.015 | [-0.066, 0.096] | 0/8 |
| Middle | 0.314 | [0.233, 0.395] | 0.298 | [0.217, 0.380] | 8/8 | |
| Late | 0.413 | [0.332, 0.494] | 0.389 | [0.308, 0.470] | 8/8 | |
| Qwen | Early | 0.011 | [-0.070, 0.092] | 0.010 | [-0.071, 0.091] | 0/8 |
| Middle | 0.342 | [0.260, 0.423] | 0.321 | [0.240, 0.402] | 8/8 | |
| Frozen path | Llama | Qwen | Mistral |
|---|---|---|---|
| NQ CQ | 0.412 [0.364, 0.460] | 0.386 [0.338, 0.434] | 0.394 [0.345, 0.442] |
| CQ NQ | 0.398 [0.351, 0.445] | 0.374 [0.326, 0.422] | – |
| NQ PopQA-C | 0.405 [0.358, 0.452] | 0.381 [0.334, 0.428] | – |
| Aligned into target | 0.342 [0.294, 0.390] | 0.326 [0.278, 0.374] | – |
| Native effect retained | 83.0% | 84.5% | – |
Appendix figures & tables21 assets
Supplementary material from the paper’s appendix.
Appendix
| Equal-norm intervention | Margin shift [95% CI] | Greedy flips | Congruent retention | No-context stability |
|---|---|---|---|---|
| Signed PC1 | .412 [.364, .460] | 24/110 | 97.8% | 96.5% |
| Signed magnitude | .365 [.317, .413] | 18/110 | 78.4% | 71.2% |
| Norm-matched random | .008 [-.040, .056] | 0/110 | 97.1% | 96.0% |
| Unsigned magnitude | .005 [-.043, .053] | 0/110 | 70.1% | 65.0% |
| Randomized sign | .008 [-.040, .056] | 0/110 | 97.1% | 96.0% |
| Sign scramble | .006 [-.042, .054] | 0/110 | 97.2% | 96.1% |
| Target | Early | Middle | Late | Source-minus-target profile interaction | Holm |
|---|---|---|---|---|---|
| Source reliance | .008 | .395 | .428 | – | – |
| Sentiment | .382 | .194 | .061 | .741 | 3/10001 |
| Format | .365 | .181 | .048 | .737 | 3/10001 |
| Language | .371 | .188 | .052 | .739 | 3/10001 |
| Condition | Context share | Parametric share | Unsupported | Mixed | Context-share change [95% CI] | NLI / quality |
|---|---|---|---|---|---|---|
| Clean | .125 | .850 | .014 | .011 | – | 98.0 / – |
| Toward contextual | .510 | .465 | .013 | .012 | [.337, .433] | 96.2 / 97.1% |
| Toward parametric | .042 | .933 | .013 | .012 | [ ] | 96.0 / 96.8% |
| Block | Direction | Comparator | Advantage | 95% CI | Std. effect | Improve | Seeds | Holm |
|---|---|---|---|---|---|---|---|---|
| Early | (toward parametric) | Train-only L2 radial | 0.009 | [-0.072, 0.090] | 0.019 | 0.875 | 11/20 | 1.000 |
| Early | (toward parametric) | Isotropic random | 0.017 | [-0.065, 0.098] | 0.034 | 0.875 | 11/20 | 1.000 |
| Early | (toward parametric) | PC1-orthogonal random | 0.019 | [-0.062, 0.100] | 0.039 | 0.875 | 11/20 | 1.000 |
| Early | (toward parametric) | Shuffled-layer PC1 | 0.011 | [-0.071, 0.092] | 0.022 | 0.875 | 11/20 | 1.000 |
| Early | (toward context) | Train-only L2 radial | 0.007 | [-0.074, 0.088] | 0.015 | 0.875 | 11/20 | 1.000 |
| Early | (toward context) | Isotropic random | 0.013 | [-0.068, 0.095] | 0.028 | 0.875 | 11/20 | 1.000 |
| Block | Direction | Comparator | Advantage | 95% CI | Std. effect | Improve | Seeds | Holm |
|---|---|---|---|---|---|---|---|---|
| Early | (toward parametric) | Train-only L2 radial | 0.003 | [-0.078, 0.084] | 0.006 | 0.874 | 10/20 | 1.000 |
| Early | (toward parametric) | Isotropic random | 0.010 | [-0.071, 0.091] | 0.021 | 0.874 | 10/20 | 1.000 |
| Early | (toward parametric) | PC1-orthogonal random | 0.011 | [-0.070, 0.093] | 0.024 | 0.874 | 10/20 | 1.000 |
| Early | (toward parametric) | Shuffled-layer PC1 | 0.005 | [-0.076, 0.086] | 0.011 | 0.874 | 10/20 | 1.000 |
| Early | (toward context) | Train-only L2 radial | 0.002 | [-0.079, 0.084] | 0.005 | 0.874 | 10/20 | 1.000 |
| Early | (toward context) | Isotropic random | 0.009 | [-0.072, 0.090] | 0.018 | 0.874 | 10/20 | 1.000 |
| Model | Block | Direction | Candidate effect [95% CI] | LTS–candidate [95% CI] | ||
|---|---|---|---|---|---|---|
| Llama | Early | C P | 0.0024 [0.0021, 0.0027] | 0.0006 | 0.0022 [0.0017, 0.0026] | 0.0006 |
| Llama | Early | P C | 0.0017 [0.0012, 0.0022] | 0.0006 | 0.0037 [0.0031, 0.0043] | 0.0006 |
| Llama | Middle | C P | 0.1407 [0.1397, 0.1417] | 0.0006 | 0.2536 [0.2524, 0.2548] | 0.0006 |
| Llama | Middle | P C | 0.1423 [0.1411, 0.1435] | 0.0006 | 0.2530 [0.2513, 0.2546] | 0.0006 |
| Llama | Late | C P | 0.3912 [0.3899, 0.3924] | 0.0006 | 0.0365 [0.0349, 0.0380] | 0.0006 |
| Llama | Late | P C | 0.3910 [0.3899, 0.3920] | 0.0006 | 0.0363 [0.0349, 0.0377] | 0.0006 |
| Model | Block | Direction | Candidate effect [95% CI] | LTS–candidate [95% CI] | ||
|---|---|---|---|---|---|---|
| Llama | Early | C P | 0.0017 [0.0008, 0.0026] | 0.0014 | 0.0030 [0.0021, 0.0041] | 0.0006 |
| Llama | Early | P C | 0.0016 [0.0009, 0.0023] | 0.0012 | 0.0034 [0.0024, 0.0045] | 0.0006 |
| Llama | Middle | C P | 0.1294 [0.1271, 0.1316] | 0.0006 | 0.2810 [0.2783, 0.2837] | 0.0006 |
| Llama | Middle | P C | 0.1294 [0.1268, 0.1321] | 0.0006 | 0.2835 [0.2797, 0.2871] | 0.0006 |
| Llama | Late | C P | 0.2139 [0.2118, 0.2159] | 0.0006 | 0.2074 [0.2046, 0.2102] | 0.0006 |
| Llama | Late | P C | 0.2142 [0.2117, 0.2168] | 0.0006 | 0.2064 [0.2032, 0.2097] | 0.0006 |
| Model | Block | Margin-eligible | Candidate flips | Flip rate | Top-1 remains A/B |
|---|---|---|---|---|---|
| Llama | Middle | 289 | 123 | 42.6% | 777/777 |
| Llama | Late | 289 | 179 | 61.9% | 777/777 |
| Qwen | Middle | 235 | 94 | 40.0% | 617/617 |
| Qwen | Late | 235 | 134 | 57.0% | 617/617 |
| Model | Endpoint | Successes/ | Point estimate | 95% CI | Cells represented | 95% gate |
|---|---|---|---|---|---|---|
| Llama | Congruent candidate accuracy | 712/777 | 91.63% | [89.48, 93.38]% | 3 blocks 2 signs | 0/6 |
| Llama | No-context choice stability | 699/777 | 89.96% | [87.65, 91.88]% | 3 blocks 2 signs | 0/6 |
| Qwen | Congruent candidate accuracy | 565/617 | 91.57% | [89.11, 93.52]% | 3 blocks 2 signs | 0/6 |
| Qwen | No-context choice stability | 555/617 | 89.95% | [87.33, 92.08]% | 3 blocks 2 signs | 0/6 |
| Model | Method | Prevalence | AP | 95% CI | LTS method | 95% CI | Holm |
|---|---|---|---|---|---|---|---|
| Llama | LTS | 0.055 | 0.311 | [0.206, 0.444] | – | – | – |
| Llama | Supervised LTS | 0.055 | 0.415 | [0.301, 0.546] | -0.103 | [-0.239, 0.037] | 0.153 |
| Llama | L2 magnitude | 0.055 | 0.741 | [0.615, 0.862] | -0.430 | [-0.596, -0.247] | |
| Llama | Answer margin | 0.055 | 0.989 | [0.973, 0.998] | -0.678 | [-0.782, -0.545] | |
| Qwen | LTS | 0.066 | 0.716 | [0.586, 0.828] | – | – | – |
| Qwen | Supervised LTS | 0.066 | 0.631 | [0.502, 0.748] | 0.085 | [-0.010, 0.185] | 0.225 |
| Model | Block | Direction | Target success [95% CI] | Comparator risk-difference range | max | Controls | |
|---|---|---|---|---|---|---|---|
| Llama | Early | 0.011 [ , 0.046] | [0.000, 0.012] | 1.0000 | 0/4 | 779 | |
| Early | 0.012 [ , 0.047] | [ , 0.004] | 1.0000 | 0/4 | 779 | ||
| Middle | 0.254 [0.219, 0.289] | [0.187, 0.192] | 0.0024 | 4/4 | 779 | ||
| Middle | 0.243 [0.208, 0.278] | [0.191, 0.194] | 0.0024 | 4/4 | 779 | ||
| Late | 0.227 [0.192, 0.262] | [0.186, 0.191] | 0.0024 | 4/4 | 779 | ||
| Late | 0.240 [0.205, 0.275] | [0.187, 0.194] | 0.0024 | 4/4 | 779 |
| Endpoint | Benchmark | Model | Observed | 95% interval | Gate | Pass |
|---|---|---|---|---|---|---|
| Eligible P C top-1 flips | NQSwap | Llama | 24/110 (21.8%) | [15.1, 30.4]% | yes | |
| NQSwap | Qwen | 20/100 (20.0%) | [13.3, 28.9]% | yes | ||
| Congruent candidate accuracy | ConflictQA | Llama | 97.8% | [96.5, 98.6]% | yes | |
| ConflictQA | Qwen | 97.8% | [96.3, 98.7]% | yes | ||
| No-context clean-choice stability | ConflictQA | Llama | 96.5% | [95.0, 97.6]% | yes | |
| ConflictQA | Qwen | 96.5% | [94.7, 97.7]% | yes |
| Block | Dose | Comparator | Llama | Qwen | ||||||
|---|---|---|---|---|---|---|---|---|---|---|
| Risk diff. [95% CI] | Gate | Risk diff. [95% CI] | Gate | |||||||
| Early | L2 radial | 0.0102 [ , 0.0402] | 1.0000 | fail | 779 | 0.0066 [ , 0.0366] | 1.0000 | fail | 617 | |
| isotropic random | 0.0004 [ , 0.0304] | 1.0000 | fail | 779 | 0.0122 [ , 0.0422] | 1.0000 | fail | 617 | ||
| PC1-orthogonal | 0.0037 [ , 0.0337] | 1.0000 | fail | 779 | 0.0094 [ , 0.0394] | 1.0000 | fail | 617 | ||
| shuffled layer | 0.0117 [ , 0.0417] | 1.0000 | fail | 779 | 0.0053 [ , 0.0353] | 1.0000 | fail | 617 | ||
| L2 radial | [ , 0.0256] | 1.0000 | fail | 779 | [ , 0.0299] | 1.0000 | fail | 617 | ||
| Block | Move | Comparator | Llama | Qwen | Mistral |
|---|---|---|---|---|---|
| Early | P C | random direction | -0.002/1.0000 | -0.002/1.0000 | -0.002/1.0000 |
| Early | P C | zero vector | 0.000/1.0000 | 0.000/1.0000 | 0.000/1.0000 |
| Early | P C | orthogonal PC2 | 0.002/1.0000 | 0.002/1.0000 | 0.002/1.0000 |
| Early | P C | layer reversed | 0.004/1.0000 | 0.004/1.0000 | 0.004/1.0000 |
| Early | C P | random direction | -0.004/1.0000 | -0.004/1.0000 | -0.004/1.0000 |
| Early | C P | zero vector | -0.002/1.0000 | -0.002/1.0000 | -0.002/1.0000 |
| Source | Target | Model | Items | Clusters | Effect [95% CI] | |
|---|---|---|---|---|---|---|
| NQSwap | ConflictQA | Llama | 777 | 50 | 0.412 [0.364, 0.460] | 0.0014 |
| NQSwap | ConflictQA | Qwen | 617 | 50 | 0.386 [0.338, 0.434] | 0.0014 |
| NQSwap | ConflictQA | Mistral | 700 | 45 | 0.394 [0.345, 0.442] | 0.0014 |
| ConflictQA | NQSwap | Llama | 120 | 25 | 0.398 [0.351, 0.445] | 0.0014 |
| ConflictQA | NQSwap | Qwen | 87 | 20 | 0.374 [0.326, 0.422] | 0.0014 |
| NQSwap | PopQA-Conflict | Llama | 500 | 40 | 0.405 [0.358, 0.452] | 0.0014 |
| Transport | Arm | Effect | 95% CI | |
|---|---|---|---|---|
| Llama Qwen | aligned LTS | 0.326 | [0.278, 0.374] | 0.0008 |
| Llama Qwen | random map | 0.012 | [-0.035, 0.059] | 1.0000 |
| Llama Qwen | wrong layer | -0.005 | [-0.052, 0.042] | 1.0000 |
| Llama Qwen | orthogonal direction | 0.008 | [-0.039, 0.055] | 1.0000 |
| Qwen Llama | aligned LTS | 0.342 | [0.294, 0.390] | 0.0008 |
| Qwen Llama | random map | 0.009 | [-0.038, 0.056] | 1.0000 |
| Model | Comparator | LTS | Comparator | LTS comparator | Holm | |
|---|---|---|---|---|---|---|
| Pythia-70M | L2 | 0.505 [0.492, 0.518] | 0.521 [0.502, 0.540] | -0.015 [-0.035, +0.005] | 0.169 | 0.634 |
| Pythia-70M | likelihood | 0.505 [0.492, 0.518] | 0.491 [0.476, 0.505] | +0.014 [-0.005, +0.034] | 0.158 | 0.634 |
| Pythia-70M | word TF–IDF | 0.505 [0.492, 0.518] | 0.491 [0.477, 0.504] | +0.015 [-0.004, +0.033] | 0.097 | 0.487 |
| Pythia-70M | char TF–IDF | 0.505 [0.492, 0.518] | 0.488 [0.475, 0.500] | +0.018 [-0.000, +0.036] | 0.042 | 0.293 |
| Pythia-70M | joint TF–IDF | 0.505 [0.492, 0.518] | 0.488 [0.475, 0.501] | +0.017 [-0.001, +0.035] | 0.052 | 0.310 |
| Pythia-70M | metadata | 0.505 [0.492, 0.518] | 0.497 [0.483, 0.512] | +0.008 [-0.011, +0.027] | 0.406 | 0.813 |
| Model | Contrast | LTS | L2 | LTS L2 | Holm | |
|---|---|---|---|---|---|---|
| Pythia-70M | relevant no context | 0.505 [0.489, 0.521] | 0.512 [0.495, 0.529] | -0.007 [-0.027, +0.013] | 0.503 | 0.503 |
| Pythia-70M | relevant shuffled | 0.500 [0.485, 0.516] | 0.488 [0.473, 0.502] | +0.013 [-0.008, +0.033] | 0.230 | 0.230 |
| Pythia-160M | relevant no context | 0.513 [0.496, 0.529] | 0.516 [0.500, 0.532] | -0.003 [-0.023, +0.016] | 0.717 | 0.717 |
| Pythia-160M | relevant shuffled | 0.499 [0.484, 0.515] | 0.502 [0.486, 0.517] | -0.002 [-0.025, +0.020] | 0.858 | 0.858 |
| Pythia-410M | relevant no context | 0.495 [0.480, 0.510] | 0.491 [0.477, 0.506] | +0.004 [-0.015, +0.023] | 0.676 | 0.676 |
| Pythia-410M | relevant shuffled | 0.512 [0.496, 0.528] | 0.514 [0.497, 0.531] | -0.002 [-0.024, +0.019] | 0.830 | 0.830 |
| Model | Best likelihood | LTS–LR [95% CI] | LTS–XGB [95% CI] | Gain over likelihood | LTS dim. |
|---|---|---|---|---|---|
| Llama-3.1-8B | 0.565 | 0.778 [0.741, 0.838] | 0.707 [0.655, 0.759] | +0.213 | 12 |
| Llama-3.1-8B-Inst | 0.575 | 0.708 [0.665, 0.760] | 0.627 [0.553, 0.687] | +0.133 | 12 |
| Mistral-7B-v0.3 | 0.575 | 0.869 [0.843, 0.890] | 0.815 [0.777, 0.856] | +0.294 | 12 |
| Mistral-7B-Inst | 0.596 | 0.799 [0.744, 0.845] | 0.731 [0.680, 0.769] | +0.203 | 12 |
| Qwen2.5-7B | 0.583 | 0.784 [0.765, 0.801] | 0.765 [0.739, 0.791] | +0.201 | 10 |
| Qwen2.5-7B-Inst | 0.579 | 0.869 [0.837, 0.901] | 0.815 [0.777, 0.860] | +0.290 | 10 |
| Analysis | Scope | Complete retained result | Interpretation |
|---|---|---|---|
| Model-free WikiMIA audit | Same 250 rows | Word TF–IDF 0.965 [0.945, 0.982]; word+character 0.962 [0.940, 0.980]; character 0.939 [0.912, 0.962]; length/style 0.645 [0.579, 0.710] | Strong shortcut warning |
| Magnitude baseline | Nine models | Multilayer L2 mean AUC 0.812 versus LTS–LR 0.836; LTS matches or exceeds L2 on 5/9 models | Distinct, neither universal |
| Same-topic control | Qwen-14B, Mistral-7B, Llama-8B | LTS–LR 0.921, 0.842, 0.726; changes from random pairing are | Topic does not explain all signal |
| Prompt/label controls | Same three models | Four-template standard deviations 0.019, 0.009, 0.011; ten label permutations return | Classifier and prompt checks pass |
| BookMIA | Qwen-14B, Mistral-7B, Llama-8B | Continuation LTS–LR 0.844, 0.967, 0.905; QA 0.980, 0.959, 0.969; L2 0.683, 0.823, 0.745 | External benchmark transfer |
| MIMIR negative control | Pile-Wikipedia split | LTS AUC range 0.48–0.55 | No universal membership detector |