Belief-Trajectory Energy: Measuring the Path to a Prediction
Organizations: Fudan University · University of Science and Technology of China · Shanghai Innovation Institute
Abstract
Large language models (LLMs) progressively revise their predictions across Transformer layers, yet we typically observe only the final output, discarding the trajectory through which it is formed. We introduce Belief-Trajectory Energy(BTE), a model-grounded measure that characterizes an input through the layerwise predictive revisions it induces in a model. By mapping intermediate states into a shared predictive space, BTE provides a principled measure of belief change that can be summarized as either a scalar or a structured depth profile. Theoretically, we show that local BTE corresponds to predictive revision under the Fisher-Rao geometry, while the sequence of revisions captures information beyond the initial-to-final belief change. Empirically, scalar BTE provides a model-relative signal of difficulty across diverse reasoning tasks, while richer BTE representations support human-LLM review detection and fine-grained generator attribution, reaching up to macro-AUROC and eight-way attribution accuracy. Further analysis shows that BTE develops throughout pretraining and is selectively reshaped by targeted training, demonstrating that the resulting measurement reflects what the scoring model has learned. Together, our results establish belief trajectories as a principled model-grounded signal and suggest a broader perspective in which learned models can themselves serve as instruments for characterizing the data they process. More demonstrations can be found at https://yingjiahao14.github.io/BTE-web/.
Figures & tables
Appendix figures & tables12 assets
Supplementary material from the paper’s appendix.
Appendix
| Solution source | [95% CI] | |
| Official references | 150 | 0.44 [0.29, 0.57] |
| GPT-5.5 rewrites | 150 | 0.75 [0.68, 0.81] |
| A | B | C | ||||
| Test conference-year | Llama | Qwen | Llama | Qwen | Llama | Qwen |
| ICLR 2022 | 0.8129 | 0.9314 | 0.8939 | 0.9997 | 0.9601 | 1.0000 |
| ICLR 2023 | 0.7556 | 0.9357 | 0.9441 | 0.9999 | 0.9827 | 0.9998 |
| ICLR 2024 | 0.7723 | 0.8729 | 0.8966 | 0.9865 | 0.9616 | 0.9932 |
| NeurIPS 2021 | 0.7430 | 0.9282 | 0.9444 | 1.0000 | 0.9822 | 1.0000 |
| NeurIPS 2022 | 0.7450 | 0.9231 | 0.9389 | 1.0000 | 0.9872 | 0.9999 |
| Test conference-year | GPT-4o | Gem-1.5 | Clau-3.5 | GPT-5.5 | GPT-5.6 | Gem-3.1 | Opus-5 |
| Llama-3.1-8B, setting A | |||||||
| ICLR 2022 | 0.999 | 0.999 | 0.999 | 0.937 | 0.860 | 0.792 | 0.105 |
| ICLR 2023 | 1.000 | 0.999 | 0.998 | 0.689 | 0.498 | 0.958 | 0.147 |
| ICLR 2024 | 0.990 | 0.968 | 0.939 | 0.788 | 0.738 | 0.895 | 0.088 |
| NeurIPS 2021 | 1.000 | 0.998 | 0.999 | 0.649 | 0.493 | 0.943 | 0.120 |
| NeurIPS 2022 | 1.000 | 0.998 | 0.996 | 0.639 | 0.510 | 0.934 | 0.139 |
| Generator | 95% CI | ||||
| GPT-4o | 23 | 6 | 31 | 0.433 | |
| Gemini 1.5 Pro | 38 | 2 | 20 | 0.650 | |
| Claude 3.5 Sonnet | 29 | 3 | 28 | 0.508 | |
| GPT-5.5 | 49 | 6 | 5 | 0.867 | |
| GPT-5.6 | 51 | 5 | 4 | 0.892 | |
| Gemini 3.1 Pro | 28 | 2 | 30 | 0.483 |
| Llama-3.1-8B-Instruction | Qwen-3.5-9B | ||||||
| Per-layer features | Dim. | A | B | C | A | B | C |
| Token mean | 32 | 72.1 | 82.9 | 86.4 | 82.9 | 89.5 | 91.3 |
| + std, P10/P50/P90 | 160 | 77.9 | 88.5 | 91.8 | 87.1 | 93.1 | 94.4 |
| + means in 4 position bins | 288 | 76.4 | 86.2 | 92.7 | 87.0 | 93.2 | 95.0 |
| + AC, top-10% mean | 352 | 77.1 | 87.1 | 93.2 | 86.6 | 94.0 | 95.6 |
| Position bins only | 128 | 70.5 | 82.1 | 88.6 | 79.6 | 90.1 | 92.6 |
| Llama-3.1-8B-Instruct | Qwen-3.5-9B | |||||
| Per-layer features | A | B | C | A | B | C |
| Token mean | 79.1 / 2.3 | 74.2 / 0.2 | 85.6 / 1.2 | 93.9 / 0.0 | 97.1 / 0.1 | 98.6 / 0.3 |
| + std, P10/P50/P90 | 85.1 / 0.7 | 94.8 / 0.5 | 93.3 / 0.8 | 94.1 / 0.0 | 97.8 / 0.1 | 98.3 / 0.1 |
| + means in 4 position bins | 86.7 / 0.8 | 94.6 / 0.5 | 93.6 / 0.5 | 94.9 / 0.1 | 98.2 / 0.1 | 98.0 / 0.0 |
| + AC, top-10% mean | 85.6 / 0.7 | 95.4 / 0.3 | 94.1 / 0.5 | 94.4 / 0.0 | 98.1 / 0.0 | 98.0 / 0.1 |
| Position bins only | 84.3 / 2.3 | 89.2 / 0.5 | 90.2 / 1.0 | 94.9 / 0.1 | 97.4 / 0.2 | 98.7 / 0.2 |
| Original | Rewrite | ||||||
| Setting | Scorer | Length | Raw 32D | Residual 32D | Length | Raw 32D | Residual 32D |
| A | Llama | 0.5383 | 0.7416 | 0.7489 | 0.5349 | 0.7197 | 0.7233 |
| A | Qwen | 0.5344 | 0.8911 | 0.8848 | 0.5308 | 0.8640 | 0.8574 |
| B | Llama | 0.5123 | 0.9094 | 0.9167 | 0.5119 | 0.9046 | 0.9127 |
| B | Qwen | 0.5344 | 0.9940 | 0.9933 | 0.5308 | 0.9934 | 0.9925 |
| C | Llama | 0.5383 | 0.9680 | 0.9695 | 0.5349 | 0.9654 | 0.9680 |
| Scorer | ||||
| Llama-3.1-8B-Instruct | 0.576 | 0.444 | 0.212 | 0.151 |
| OpenMath2-Llama3.1-8B | 0.431 | 0.391 | 0.219 | 0.191 |
| DeepSeek-R1-Distill-Llama-8B | 0.106 | 0.143 | 0.104 | 0.082 |
| Nemotron-Nano-8B-v1 | 0.219 | 0.164 | 0.061 |
| Checkpoint | Step | Tokens | L0-3 | L12-19 | L28-31 | 95% CI | |
| Gaussian init (same architecture) | - | 0 | 0.130 | 0.013 | 0.007 | ||
| stage1-step150 | 150 | 1B | 0.123 | 0.006 | 0.002 | ||
| stage1-step1000 | 1,000 | 5B | 0.139 | 0.012 | 0.007 | ||
| stage1-step5000 | 5,000 | 21B | 0.230 | 0.031 | 0.015 | ||
| stage1-step12000 | 12,000 | 51B | 0.289 | 0.044 | 0.020 | ||
| stage1-step24000 | 24,000 | 101B | 0.344 | 0.058 | 0.021 |