The Attribution Blind Spot: Layerwise Trajectory Diagnostics for Source Reliance in Retrieval-Augmented Language Models
Authors: Zhe Yu, Wenpeng Xing, Yunzhao Wei, Bo Yang, Chen Ye, Gaolei Li, Meng Han
Organizations: Zhejiang University · Binjiang Institute of Zhejiang University · East China Normal University · National FinTech Evaluation Center (Bank Card Testing Center) · Hangzhou Dianzi University · Shanghai Jiaotong University
A retrieval-augmented model can match a document without relying on it. Controlled knowledge conflicts make source choice observable and let us ask a second question that prediction alone cannot answer: which internal-state properties define useful intervention directions? We study paired hidden-state changes with Latent Trajectory Shift (LTS), a signed projection onto a training-fitted first principal component (PC1), and keep verified training exposure separate from behavioral source choice. Across the evaluated conflicts, state-change magnitude is often the stronger predictor, whereas signed PC1 is the stronger selective controller: equal-norm interventions change source preference while better preserving non-target behavior, and the frozen direction transfers across the tested datasets and aligned model pairs. A same-system OLMo study further combines positive choice and control results with inconclusive exposure detection at the achieved power. The central result is a separation: representations that diagnose what a model will choose need not be the representations that best control that choice.
Figures & tables
Figure 1: From ambiguous agreement to separate prediction and intervention tests. (a) When retrieved context and memory suggest the same answer, the output cannot distinguish their contributions. A controlled conflict makes the two answer choices distinguishable. (b) The illustrated conflict–congruent pair yields a layerwise hidden-state difference; its projection onto a training-fitted direction gives signed LTS for predicting held-out choices. (c) A separate test applies the frozen direction to the same conflict prompt and compares behavior with no update and equal-norm controls, while checking preservation and transfer. Exposure is evaluated separately in RQ1.
Arm
What it tests
Rep.
Train-only PC1
Signed source direction
1
Candidate gradient
Direct answer steering
1
L2 radial
Magnitude without PC1
1
Shuffled-layer PC1
Layer specificity
1
Isotropic random
Direction specificity
20
PC1-orthogonal random
PC1 specificity
20
Table 1: Equal-budget intervention controls. Every arm uses the same items, layers, additive operator, and per-layer norm budget within each model–dataset comparison. All use doses {−2,−1,−0.5,−0.25,0,0.25,0.5,1,2} . Rep. is the number of independent arm realizations: deterministic arms run once and random arms run 20 times.
Figure 2: Verified exposure is not detected by primary LTS in five Pythia checkpoints. (a) LTS and matched baselines across exact checkpoints. (b) Model-free text and metadata shortcuts. (c) No-context and shuffled paired controls. Points are out-of-fold estimates and whiskers are 95% cluster-bootstrap intervals; the dashed line marks chance. Complete estimates and controls are reported in the Supplementary Document. Overlap with chance is non-detection, not an equivalence claim.
Test
Endpoint
Estimate [95% CI]
Result
Exposure
ROC–AUC
0.508 [0.482, 0.534]
Inconclusive
Choice
PR–AUC
0.718 [0.667, 0.769]
Positive
Control
Δm
0.392 [0.344, 0.440]
Positive
Transfer
Δm
0.384 [0.336, 0.432]
Positive
Table 2: Four tests in one OLMo-2 system separate exposure from source choice and control. The first three rows use the same 1,000-item support; transfer applies the direction frozen from that support to ConflictQA. Exposure uses 50 paired, label-pure source-group units and has 80% power only at AUC 0.570. Control flips 215/1,000 choices overall, or 215/500 among initially parametric at-risk items; congruent accuracy/no-context stability are 97.4/96.2%.
Context condition
Parametric (%)
Contextual (%)
Other (%)
Congruent
99.1
0.0
0.9
Conflict
31.0
63.7
5.3
Unrelated, matched
88.5
7.1
4.4
Shuffled conflict
60.2
31.9
8.0
Paraphrased conflict
34.5
64.6
0.9
Mixed A/B
64.6
7.1
28.3
Table 3: Source choices under NQSwap controls. Llama-3.1-8B-Instruct, n=113 ; all outcomes retained.
Model / items
Method
ROC–AUC ↑
AP ↑
AP 95% CI
Llama n=777
LTS
0.859
0.311
[0.206, 0.444]
Supervised LTS
0.880
0.415
[0.301, 0.546]
L2 magnitude
0.978
0.741
[0.615, 0.862]
Answer margin
0.999
0.989
[0.973, 0.998]
Qwen n=617
LTS
0.903
0.716
[0.586, 0.828]
Supervised LTS
0.912
0.631
[0.502, 0.748]
Table 4: ConflictQA source-choice prediction. Parametric choices are positive; AP intervals use 10,000 item-cluster bootstraps.
Model / items
Block
Toward parametric
Toward contextual
Controls
Effect
95% CI
Effect
95% CI
passed
Llama n=120
Early
0.018
[-0.063, 0.100]
0.015
[-0.066, 0.096]
0/8
Middle
0.314
[0.233, 0.395]
0.298
[0.217, 0.380]
8/8
Late
0.413
[0.332, 0.494]
0.389
[0.308, 0.470]
8/8
Qwen n=87
Early
0.011
[-0.070, 0.092]
0.010
[-0.071, 0.091]
0/8
Middle
0.342
[0.260, 0.423]
0.321
[0.240, 0.402]
8/8
Table 5: Bidirectional source-margin effects on NQSwap. Positive effects follow the target direction. Counts cover four matched controls in both directions.
Figure 3: Prediction and control favor different state properties. (a) ConflictQA source-choice ranking for Llama ( n=777 ) and Qwen ( n=617 ), with parametric choice as the positive class. Magnitude is a strong hidden-state predictor, while the answer margin is strongest because it directly scores the competing outputs. (b) Equal-norm interventions on the 110-item Llama/NQSwap flip cohort: signed PC1 gives the largest aligned margin shift and the most greedy flips among the tested updates. Whiskers in (a–b) show 95% bootstrap intervals. (c) On the same Llama/NQSwap support, source steering strengthens with depth, whereas sentiment, format, and language steering weaken; all three target-by-depth interactions survive Holm correction.
Frozen path
Llama
Qwen
Mistral
NQ → CQ
0.412 [0.364, 0.460]
0.386 [0.338, 0.434]
0.394 [0.345, 0.442]
CQ → NQ
0.398 [0.351, 0.445]
0.374 [0.326, 0.422]
–
NQ → PopQA-C
0.405 [0.358, 0.452]
0.381 [0.334, 0.428]
–
Aligned into target
0.342 [0.294, 0.390]
0.326 [0.278, 0.374]
–
Native effect retained
83.0%
84.5%
–
Table 6: Cross-architecture and cross-dataset transfer. Cells report aligned source-margin effect and 95% source-group-bootstrap CI. NQ/CQ denote NQSwap/ConflictQA. All seven dataset paths have Holm p=0.0014 ; aligned transport has Holm p=0.0008 in each four-arm directional family. Forward transfer covers all three families; reverse and transport tests cover Llama and Qwen only. Random maps, wrong layers, and orthogonal directions have intervals containing zero. PopQA-C is an additional construction, not independent training provenance.
Appendix figures & tables21 assets
Supplementary material from the paper’s appendix.
Appendix
Equal-norm intervention
Margin shift [95% CI]
Greedy flips
Congruent retention
No-context stability
Signed PC1
.412 [.364, .460]
24/110
97.8%
96.5%
Signed ℓ2 magnitude
.365 [.317, .413]
18/110
78.4%
71.2%
Norm-matched random
.008 [-.040, .056]
0/110
97.1%
96.0%
Unsigned ℓ2 magnitude
.005 [-.043, .053]
0/110
70.1%
65.0%
Randomized sign
.008 [-.040, .056]
0/110
97.1%
96.0%
Sign scramble
.006 [-.042, .054]
0/110
97.2%
96.1%
Appendix
Table 9: Seven intervention arms separate sign from magnitude on the same 110-item Llama/NQSwap support. Signed ℓ2 applies the LTS sign to a radial update under the shared ℓ2 budget. All arms use the same items and norm budget; the maximum observed relative norm error is 4.2×10−7 . Across the five prespecified PC1 comparisons, positive margin and flip contrasts have Holm p=5/10001 . Ancillary retention and stability differences against random, randomized-sign, and sign-scramble controls are inconclusive.
Target
Early
Middle
Late
Source-minus-target profile interaction
Holm p
Source reliance
.008
.395
.428
–
–
Sentiment
.382
.194
.061
.741
3/10001
Format
.365
.181
.048
.737
3/10001
Language
.371
.188
.052
.739
3/10001
Appendix
Table 10: Source control has a different depth profile from three unrelated steering targets on the same 110-item Llama/NQSwap support. Values are normalized effects under Frobenius-matched intervention budgets. Each interaction compares the raw three-block source profile with one unrelated target; all raw p -values are 1/10001 . The observed norm error is 4.2×10−7 , below the 10−6 tolerance.
Condition
Context share
Parametric share
Unsupported
Mixed
Context-share change [95% CI]
NLI / quality
Clean
.125
.850
.014
.011
–
98.0 / –
Toward contextual
.510
.465
.013
.012
.385 [.337, .433]
96.2 / 97.1%
Toward parametric
.042
.933
.013
.012
−.083 [ −.131,−.035 ]
96.0 / 96.8%
Appendix
Table 11: Long-form, paragraph-level source endpoints for OLMo-2-1124-7B-Instruct on the preregistered long-form conflict set. Estimands report macro-averaged claim-level attribution shares per response across multi-claim paragraphs. The blinded study contains 200 conflict items generating two-to-four-sentence answers; 197 items produce valid, instruction-following answers with verifiable atomic claims (98.5%, Wilson 95% CI [0.9568, 0.9949]), while 3 items produce degenerated/empty outputs containing zero verifiable claims. The macro-averaged claim shares are evaluated on the identical paired cohort of 197 valid responses across all three conditions ( N=197 ; the 3 excluded items produced degenerated zero-claim outputs across all conditions, ensuring no paired differences arise from cohort composition changes). Across these 197 paired responses, there are approximately 780 atomic claims in each experimental condition (averaging ≈4.0 claims per valid response, 780/197≈3.96 ; specifically 780 claims in the clean baseline, 782 under contextual steering, and 778 under parametric steering). Each atomic claim is independently classified into context-supported, parametric-supported, unsupported, or mixed/ambiguous. Because each valid response’s individual claim shares sum to 1.000, the macro-averaged shares strictly sum to 1.000 in every row ( 0.125+0.850+0.014+0.011=1.000 ; 0.510+0.465+0.013+0.012=1.000 ; 0.042+0.933+0.013+0.012=1.000 ). The two directional changes have Holm p=2/10001 . Inter-annotator agreement is κ=0.88 , and human–automatic agreement is κ=0.86 . The quality noninferiority interval is [−0.018,0.024] . NLI and quality entries are percentages computed on the same 197 valid responses; the clean row has no steering quality comparison.
Block
Direction
Comparator
Advantage
95% CI
Std. effect
Improve
Seeds
Holm p
Early
+PC1 (toward parametric)
Train-only L2 radial
0.009
[-0.072, 0.090]
0.019
0.875
11/20
1.000
Early
+PC1 (toward parametric)
Isotropic random
0.017
[-0.065, 0.098]
0.034
0.875
11/20
1.000
Early
+PC1 (toward parametric)
PC1-orthogonal random
0.019
[-0.062, 0.100]
0.039
0.875
11/20
1.000
Early
+PC1 (toward parametric)
Shuffled-layer PC1
0.011
[-0.071, 0.092]
0.022
0.875
11/20
1.000
Early
−PC1 (toward context)
Train-only L2 radial
0.007
[-0.074, 0.088]
0.015
0.875
11/20
1.000
Early
−PC1 (toward context)
Isotropic random
0.013
[-0.068, 0.095]
0.028
0.875
11/20
1.000
Appendix
Table 13: Complete Llama-3.1-8B-Instruct PC1-versus-control family. Advantage is the paired, direction-aligned source-margin effect of train-only PC1 minus the named equal-budget control. No comparator, direction, or block is omitted. Holm correction is applied over all 24 cells in this model.
Block
Direction
Comparator
Advantage
95% CI
Std. effect
Improve
Seeds
Holm p
Early
+PC1 (toward parametric)
Train-only L2 radial
0.003
[-0.078, 0.084]
0.006
0.874
10/20
1.000
Early
+PC1 (toward parametric)
Isotropic random
0.010
[-0.071, 0.091]
0.021
0.874
10/20
1.000
Early
+PC1 (toward parametric)
PC1-orthogonal random
0.011
[-0.070, 0.093]
0.024
0.874
10/20
1.000
Early
+PC1 (toward parametric)
Shuffled-layer PC1
0.005
[-0.076, 0.086]
0.011
0.874
10/20
1.000
Early
−PC1 (toward context)
Train-only L2 radial
0.002
[-0.079, 0.084]
0.005
0.874
10/20
1.000
Early
−PC1 (toward context)
Isotropic random
0.009
[-0.072, 0.090]
0.018
0.874
10/20
1.000
Appendix
Table 14: Complete Qwen2.5-7B-Instruct PC1-versus-control family. Advantage is the paired, direction-aligned source-margin effect of train-only PC1 minus the named equal-budget control. No comparator, direction, or block is omitted. Holm correction is applied over all 24 cells in this model.
Figure 4: Complete dose response for both models, all three blocks, and all five intervention arms. Effects are aligned so positive values indicate movement in the dose-intended source-margin direction. Shaded bands are source-group cluster-bootstrap 95% intervals. The exact zero-dose rows are identically zero by design.
Model
Block
Direction
Candidate effect [95% CI]
pH
LTS–candidate [95% CI]
pH
Llama
Early
C → P
0.0024 [0.0021, 0.0027]
0.0006
0.0022 [0.0017, 0.0026]
0.0006
Llama
Early
P → C
0.0017 [0.0012, 0.0022]
0.0006
0.0037 [0.0031, 0.0043]
0.0006
Llama
Middle
C → P
0.1407 [0.1397, 0.1417]
0.0006
0.2536 [0.2524, 0.2548]
0.0006
Llama
Middle
P → C
0.1423 [0.1411, 0.1435]
0.0006
0.2530 [0.2513, 0.2546]
0.0006
Llama
Late
C → P
0.3912 [0.3899, 0.3924]
0.0006
0.0365 [0.0349, 0.0380]
0.0006
Llama
Late
P → C
0.3910 [0.3899, 0.3920]
0.0006
0.0363 [0.0349, 0.0377]
0.0006
Appendix
Table 15: Complete within-dataset candidate-gradient comparison. Effects are direction-aligned source-margin changes. LTS–candidate is paired within the same intervention cell. Every row passes the prespecified positive effect, positive lower bound, and Holm- p<0.05 gate; no block or direction is omitted. Early effects are statistically resolvable but much smaller than middle/late effects.
Model
Block
Direction
Candidate effect [95% CI]
pH
LTS–candidate [95% CI]
pH
Llama
Early
C → P
0.0017 [0.0008, 0.0026]
0.0014
0.0030 [0.0021, 0.0041]
0.0006
Llama
Early
P → C
0.0016 [0.0009, 0.0023]
0.0012
0.0034 [0.0024, 0.0045]
0.0006
Llama
Middle
C → P
0.1294 [0.1271, 0.1316]
0.0006
0.2810 [0.2783, 0.2837]
0.0006
Llama
Middle
P → C
0.1294 [0.1268, 0.1321]
0.0006
0.2835 [0.2797, 0.2871]
0.0006
Llama
Late
C → P
0.2139 [0.2118, 0.2159]
0.0006
0.2074 [0.2046, 0.2102]
0.0006
Llama
Late
P → C
0.2142 [0.2117, 0.2168]
0.0006
0.2064 [0.2032, 0.2097]
0.0006
Appendix
Table 16: Complete zero-refit candidate-gradient comparison. NQSwap policies are applied to ConflictQA without target refitting. All four middle/late LTS-advantage cells pass in each model. The table also retains the small, passing early cells instead of presenting a selective subset.
Model
Block
Margin-eligible n
Candidate flips
Flip rate
Top-1 remains A/B
Llama
Middle
289
123
42.6%
777/777
Llama
Late
289
179
61.9%
777/777
Qwen
Middle
235
94
40.0%
617/617
Qwen
Late
235
134
57.0%
617/617
Appendix
Table 17: Candidate-gradient discrete endpoint under zero-refit transfer. Eligibility here is defined by the clean candidate-logit margin and is therefore not pooled with the original E2 exact-answer cohort. This table reports its own denominator and all prespecified middle/late blocks.
Model
Endpoint
Successes/ n
Point estimate
95% CI
Cells represented
95% gate
Llama
Congruent candidate accuracy
712/777
91.63%
[89.48, 93.38]%
3 blocks × 2 signs
0/6
Llama
No-context choice stability
699/777
89.96%
[87.65, 91.88]%
3 blocks × 2 signs
0/6
Qwen
Congruent candidate accuracy
565/617
91.57%
[89.11, 93.52]%
3 blocks × 2 signs
0/6
Qwen
No-context choice stability
555/617
89.95%
[87.33, 92.08]%
3 blocks × 2 signs
0/6
Appendix
Table 18: Complete candidate-gradient preservation endpoints. Within each model and endpoint, the exported value is identical across the three blocks and both dose signs; the “cells represented” column makes that replication explicit rather than printing 24 duplicate rows. Gate decisions use point estimates only.
Model
Method
Prevalence
AP
95% CI
LTS − method
95% CI
Holm p
Llama
LTS
0.055
0.311
[0.206, 0.444]
–
–
–
Llama
Supervised LTS
0.055
0.415
[0.301, 0.546]
-0.103
[-0.239, 0.037]
0.153
Llama
L2 magnitude
0.055
0.741
[0.615, 0.862]
-0.430
[-0.596, -0.247]
<0.001
Llama
Answer margin
0.055
0.989
[0.973, 0.998]
-0.678
[-0.782, -0.545]
<0.001
Qwen
LTS
0.066
0.716
[0.586, 0.828]
–
–
–
Qwen
Supervised LTS
0.066
0.631
[0.502, 0.748]
0.085
[-0.010, 0.185]
0.225
Appendix
Table 19: Imbalance-aware ConflictQA source-choice ranking. AP is averaged over 20 out-of-fold split seeds. Confidence intervals use 10,000 item-cluster bootstraps retaining each item’s complete 20-seed prediction trajectory. Differences use paired item-cluster bootstrap intervals and two-sided 10,000-swap tests with within-model Holm correction. Positive differences favor LTS. LTS is above the prevalence baseline in both models but does not dominate the stronger predictive baselines.
Model
Block
Direction
Target success [95% CI]
Comparator risk-difference range
max pH
Controls
n
Llama
Early
+
0.011 [ −0.024 , 0.046]
[0.000, 0.012]
1.0000
0/4
779
Early
−
0.012 [ −0.023 , 0.047]
[ −0.004 , 0.004]
1.0000
0/4
779
Middle
+
0.254 [0.219, 0.289]
[0.187, 0.192]
0.0024
4/4
779
Middle
−
0.243 [0.208, 0.278]
[0.191, 0.194]
0.0024
4/4
779
Late
+
0.227 [0.192, 0.262]
[0.186, 0.191]
0.0024
4/4
779
Late
−
0.240 [0.205, 0.275]
[0.187, 0.194]
0.0024
4/4
779
Appendix
Table 20: E2 full-set absolute directional endpoints on ConflictQA. Absolute intervals use 10,000 item bootstraps. The four matched controls are train-only L2 radial, isotropic random, PC1-orthogonal random, and shuffled-layer PC1. + targets parametric answer A , − targets contextual answer B , and target success is aligned to the indicated direction. Comparator ranges and maximum Holm p summarize 48 cells across both models (24 per model). Early blocks are the prespecified failure region.
Endpoint
Benchmark
Model
Observed
95% interval
Gate
Pass
Eligible P → C top-1 flips
NQSwap
Llama
24/110 (21.8%)
[15.1, 30.4]%
>0
yes
NQSwap
Qwen
20/100 (20.0%)
[13.3, 28.9]%
>0
yes
Congruent candidate accuracy
ConflictQA
Llama
97.8%
[96.5, 98.6]%
≥95%
yes
ConflictQA
Qwen
97.8%
[96.3, 98.7]%
≥95%
yes
No-context clean-choice stability
ConflictQA
Llama
96.5%
[95.0, 97.6]%
≥95%
yes
ConflictQA
Qwen
96.5%
[94.7, 97.7]%
≥95%
yes
Appendix
Table 21: Eligible discrete flips and preservation point-estimate gates across benchmarks. Discrete top-1 flip rates are evaluated on prespecified independent NQSwap diagnostic cohorts (110 items for Llama, 100 items for Qwen), where each item’s unsteered baseline choice under the conflicting prompt is parametric, measuring flips to contextual ( P→C ) under the identical conflicting prompt. Wilson intervals are reported on these explicit eligible subsets. Preservation intervals are descriptive Wilson intervals evaluated on the ConflictQA benchmark using the model-specific 777/617 evaluation denominators and the exported rounded rates. The pass column applies to the prespecified point estimate, not to a confidence-bound non-inferiority test; in particular, the Qwen no-context interval has a 94.7% lower bound.
Block
Dose
Comparator
Llama
Qwen
Risk diff. [95% CI]
pH
Gate
n
Risk diff. [95% CI]
pH
Gate
n
Early
+
L2 radial
0.0102 [ −0.0198 , 0.0402]
1.0000
fail
779
0.0066 [ −0.0234 , 0.0366]
1.0000
fail
617
+
isotropic random
0.0004 [ −0.0296 , 0.0304]
1.0000
fail
779
0.0122 [ −0.0178 , 0.0422]
1.0000
fail
617
+
PC1-orthogonal
0.0037 [ −0.0263 , 0.0337]
1.0000
fail
779
0.0094 [ −0.0206 , 0.0394]
1.0000
fail
617
+
shuffled layer
0.0117 [ −0.0183 , 0.0417]
1.0000
fail
779
0.0053 [ −0.0247 , 0.0353]
1.0000
fail
617
−
L2 radial
−0.0044 [ −0.0344 , 0.0256]
1.0000
fail
779
−0.0001 [ −0.0301 , 0.0299]
1.0000
fail
617
Appendix
Table 22: Complete E2 matched-control ledger (48 cells). Every exported effect, 95% interval, within-model Holm value, gate result, and denominator is shown. All 16 early cells fail because their intervals cross zero; all 32 middle/late cells pass.
Block
Move
Comparator
Llama Δ/pH
Qwen Δ/pH
Mistral Δ/pH
Early
P → C
random direction
-0.002/1.0000
-0.002/1.0000
-0.002/1.0000
Early
P → C
zero vector
0.000/1.0000
0.000/1.0000
0.000/1.0000
Early
P → C
orthogonal PC2
0.002/1.0000
0.002/1.0000
0.002/1.0000
Early
P → C
layer reversed
0.004/1.0000
0.004/1.0000
0.004/1.0000
Early
C → P
random direction
-0.004/1.0000
-0.004/1.0000
-0.004/1.0000
Early
C → P
zero vector
-0.002/1.0000
-0.002/1.0000
-0.002/1.0000
Appendix
Table 24: Complete zero-refit forward-transfer family (72 cells). Each entry is the aligned LTS-minus-comparator effect and its within-model 24-cell Holm value. Tests use 10,000 source-group sign-flip permutations with Laplace add-one correction. All 48 middle/late cells pass the prespecified positive-effect and pH<0.05 gate; all 24 early cells fail. No cell is omitted.
Source
Target
Model
Items
Clusters
Effect [95% CI]
pH
NQSwap
ConflictQA
Llama
777
50
0.412 [0.364, 0.460]
0.0014
NQSwap
ConflictQA
Qwen
617
50
0.386 [0.338, 0.434]
0.0014
NQSwap
ConflictQA
Mistral
700
45
0.394 [0.345, 0.442]
0.0014
ConflictQA
NQSwap
Llama
120
25
0.398 [0.351, 0.445]
0.0014
ConflictQA
NQSwap
Qwen
87
20
0.374 [0.326, 0.422]
0.0014
NQSwap
PopQA-Conflict
Llama
500
40
0.405 [0.358, 0.452]
0.0014
Appendix
Table 25: Complete path-level zero-refit transfer family. Intervals use 10,000 source-group bootstraps. Each two-sided sign-flip test retains one extreme draw, so praw=2/10001 ; Holm adjustment over all seven prespecified paths gives pH=0.0014 . PopQA-Conflict is an additional construction rather than independent training provenance.
Transport
Arm
Effect
95% CI
pH
Llama → Qwen
aligned LTS
0.326
[0.278, 0.374]
0.0008
Llama → Qwen
random map
0.012
[-0.035, 0.059]
1.0000
Llama → Qwen
wrong layer
-0.005
[-0.052, 0.042]
1.0000
Llama → Qwen
orthogonal direction
0.008
[-0.039, 0.055]
1.0000
Qwen → Llama
aligned LTS
0.342
[0.294, 0.390]
0.0008
Qwen → Llama
random map
0.009
[-0.038, 0.056]
1.0000
Appendix
Table 26: Complete bidirectional cross-model transport family. The learned maps retain 84.5% of the native Qwen effect and 83.0% of the native Llama effect. Intervals use 10,000 source-group bootstraps. In each four-arm family, the aligned arm has pH=0.0008 ; all three controls include zero and have pH=1 .
Model
Comparator
LTS
Comparator
LTS − comparator
p
Holm p
Pythia-70M
L2
0.505 [0.492, 0.518]
0.521 [0.502, 0.540]
-0.015 [-0.035, +0.005]
0.169
0.634
Pythia-70M
likelihood
0.505 [0.492, 0.518]
0.491 [0.476, 0.505]
+0.014 [-0.005, +0.034]
0.158
0.634
Pythia-70M
word TF–IDF
0.505 [0.492, 0.518]
0.491 [0.477, 0.504]
+0.015 [-0.004, +0.033]
0.097
0.487
Pythia-70M
char TF–IDF
0.505 [0.492, 0.518]
0.488 [0.475, 0.500]
+0.018 [-0.000, +0.036]
0.042
0.293
Pythia-70M
joint TF–IDF
0.505 [0.492, 0.518]
0.488 [0.475, 0.501]
+0.017 [-0.001, +0.035]
0.052
0.310
Pythia-70M
metadata
0.505 [0.492, 0.518]
0.497 [0.483, 0.512]
+0.008 [-0.011, +0.027]
0.406
0.813
Appendix
Table 28: Complete fold-aligned primary comparisons. Differences use identical examples, groups, folds, and split seeds; intervals use 10,000 paired group bootstraps and p -values use 10,000 paired group permutations with Holm correction within each model.
Model
Contrast
LTS
L2
LTS − L2
p
Holm p
Pythia-70M
relevant − no context
0.505 [0.489, 0.521]
0.512 [0.495, 0.529]
-0.007 [-0.027, +0.013]
0.503
0.503
Pythia-70M
relevant − shuffled
0.500 [0.485, 0.516]
0.488 [0.473, 0.502]
+0.013 [-0.008, +0.033]
0.230
0.230
Pythia-160M
relevant − no context
0.513 [0.496, 0.529]
0.516 [0.500, 0.532]
-0.003 [-0.023, +0.016]
0.717
0.717
Pythia-160M
relevant − shuffled
0.499 [0.484, 0.515]
0.502 [0.486, 0.517]
-0.002 [-0.025, +0.020]
0.858
0.858
Pythia-410M
relevant − no context
0.495 [0.480, 0.510]
0.491 [0.477, 0.506]
+0.004 [-0.015, +0.023]
0.676
0.676
Pythia-410M
relevant − shuffled
0.512 [0.496, 0.528]
0.514 [0.497, 0.531]
-0.002 [-0.024, +0.019]
0.830
0.830
Appendix
Table 29: Complete exact-checkpoint structural controls. Values are pooled out-of-fold ROC–AUC [95% CI]. LTS–L2 differences use 10,000 paired group bootstraps and permutation tests over the same 3,994 probes/1,997 groups and 20 split seeds. Each prespecified model–condition LTS–L2 contrast is a singleton family, so its displayed Holm and raw p -values coincide.
Model
Best likelihood
LTS–LR [95% CI]
LTS–XGB [95% CI]
Gain over likelihood
LTS dim.
Llama-3.1-8B
0.565
0.778 [0.741, 0.838]
0.707 [0.655, 0.759]
+0.213
12
Llama-3.1-8B-Inst
0.575
0.708 [0.665, 0.760]
0.627 [0.553, 0.687]
+0.133
12
Mistral-7B-v0.3
0.575
0.869 [0.843, 0.890]
0.815 [0.777, 0.856]
+0.294
12
Mistral-7B-Inst
0.596
0.799 [0.744, 0.845]
0.731 [0.680, 0.769]
+0.203
12
Qwen2.5-7B
0.583
0.784 [0.765, 0.801]
0.765 [0.739, 0.791]
+0.201
10
Qwen2.5-7B-Inst
0.579
0.869 [0.837, 0.901]
0.815 [0.777, 0.860]
+0.290
10
Appendix
Table 30: Complete original nine-model WikiMIA discovery screen. Best likelihood is the maximum of perplexity, zlib-normalized perplexity, and Min- K% probability. Brackets are bootstrap intervals. The temporal label and shared-calibration protocol make this a breadth result, not verified model-relative exposure.
Analysis
Scope
Complete retained result
Interpretation
Model-free WikiMIA audit
Same 250 rows
Word TF–IDF 0.965 [0.945, 0.982]; word+character 0.962 [0.940, 0.980]; character 0.939 [0.912, 0.962]; length/style 0.645 [0.579, 0.710]
Strong shortcut warning
Magnitude baseline
Nine models
Multilayer L2 mean AUC 0.812 versus LTS–LR 0.836; LTS matches or exceeds L2 on 5/9 models
Distinct, neither universal
Same-topic control
Qwen-14B, Mistral-7B, Llama-8B
LTS–LR 0.921, 0.842, 0.726; changes from random pairing are −0.004,+0.020,−0.058
Topic does not explain all signal
Prompt/label controls
Same three models
Four-template standard deviations 0.019, 0.009, 0.011; ten label permutations return 0.50±0.05
Table 31: Compact retention of the legacy robustness and boundary analyses. These rows preserve positive, mixed, and negative outcomes. The MIMIR negative control (Duan et al. 2024; full citation in the main paper) and the model-free WikiMIA audit prevent the discovery screen from being interpreted as verified exposure or direct source use.
Retrieval-augmented generation promises to ground language model outputs in external evidence, yet the field has no reliable way to verify whether retrieved context actually governs generation -- a prerequisite for any high-stakes deployment. The standard assumption, that context-consistent output implies context-governed output, breaks when the retrieved document overlaps with the model's pretraining data: the model can produce faithful-looking text entirely from parametric memory, and both pathways yield indistinguishable output. We name this failure the attribution blind spot and introduce Computational Reality Monitoring (CRM) to address it. CRM operationalizes a principle adapted from cognitive science's reality monitoring framework: comparing internal representations with and without context reveals membership-conditioned representational divergence that output-level monitors systematically miss. CRM does not certify which source an individual generation used; it detects whether pretraining exposure leaves a measurable internal trajectory signature, establishing a necessary substrate for source attribution. Across nine model variants spanning three families, this divergence concentrates in architecture-specific layer patterns, receives converging support from block-level noise intervention, and generalizes across tasks and datasets while collapsing on domain-confounded benchmarks. The attribution blind spot is measurable and partially addressable: internal representations carry a diagnostic signal invisible at the output level, establishing a foundation for systems whose internal awareness of evidence provenance governs their external behavior.
Zhe Yu, Wenpeng Xing, Yunzhao Wei +4
Binjiang Institute of Zhejiang University · Zhejiang University · National Fintech Evaluation Center +2
Large language models (LLMs) progressively revise their predictions across Transformer layers, yet we typically observe only the final output, discarding the trajectory through which it is formed. We introduce Belief-Trajectory Energy(BTE), a model-grounded measure that characterizes an input through the layerwise predictive revisions it induces in a model. By mapping intermediate states into a shared predictive space, BTE provides a principled measure of belief change that can be summarized as either a scalar or a structured depth profile. Theoretically, we show that local BTE corresponds to predictive revision under the Fisher-Rao geometry, while the sequence of revisions captures information beyond the initial-to-final belief change. Empirically, scalar BTE provides a model-relative signal of difficulty across diverse reasoning tasks, while richer BTE representations support human-LLM review detection and fine-grained generator attribution, reaching up to 0.998 macro-AUROC and 95.6% eight-way attribution accuracy. Further analysis shows that BTE develops throughout pretraining and is selectively reshaped by targeted training, demonstrating that the resulting measurement reflects what the scoring model has learned. Together, our results establish belief trajectories as a principled model-grounded signal and suggest a broader perspective in which learned models can themselves serve as instruments for characterizing the data they process. More demonstrations can be found at https://yingjiahao14.github.io/BTE-web/.
Jiahao Ying, Wei Tang, Boxian Ai +7
Fudan University · University of Science and Technology of China · Shanghai Innovation Institute
Retrieval-augmented question answering requires control decisions about when to decompose a question, search, reformulate, extract evidence, synthesize facts, verify progress, and stop. We study whether trajectory fine-tuning can improve small language models (SLMs) as next-action controllers. We additionally evaluate a low-resource setting in which a single SLM serves as both the controller and the final-answer generator. From accepted teacher search traces, we build a seven-way action-prediction task, where the model predicts the next structured teacher action from the current trajectory state, and evaluate LoRA-supervised fine-tuning across SLMs and xSLMs as controllers. On 1,646 held-out action examples, Granite 4.1 3B trained on 13,194 actions reaches macro-F1 0.6536, compared with 0.1736 for zero-shot prompting of the same model and 0.5399 for a TF-IDF logistic-regression baseline. In an end-to-end controller/generator swap evaluation over 149 held-out trajectories, using the fine-tuned model for both roles improves Exact Match from 0.7530 to 0.7946 and token F1 from 0.7783 to 0.8295 compared with using the base model as both controller and generator. The cross-role conditions show that the fine-tuned controller increases evidence-fact recording when the generator is fixed, while controller-only final-answer gains are not statistically clear. Overall, trajectory supervision improves action prediction and evidence-recording behaviour in this evaluated pipeline. Code is available at https://github.com/padas-lab-de/agent-action-controller
Mohammed Al-Maamari, Saber Zerhoudi, Michael Granitzer +1