Reliable confidence estimation is essential for large language model deployment. However, answer-level calibration remains challenging because generation errors are often localized: a response may be fluent and high-probability overall while still failing at a critical number, entity, or factual claim. Existing estimators compress token probabilities, sequence likelihoods, entropy, or beam statistics into a global score, which can dilute such local risk signals. We propose TRACE, a single-pass, decoded-answer-preserving confidence estimator that treats decoding-time uncertainty as a trajectory through three steps: (i) recording token-level surprisal and predictive entropy during decoding, (ii) applying local risk operators to preserve uncertainty spikes, and (iii) converting localized trace risk into answer-level confidence. TRACE produces a label-free risk score, while TRACE+ calibrates trace-only features into probabilities using a held-out split, without extra generations or external verifiers. We evaluate four tasks against 19 calibration baselines, and TRACE+ reduces Brier from 0.149 to 0.137 and improves AUROC from 0.758 to 0.792 over the strongest likelihood baseline. Across seven LLMs, TRACE+ improves over the best non-TRACE baseline pool from 0.136 to 0.120 Brier and from 0.764 to 0.817 AUROC. Results show that localizing decoding-time risk provides a general approach to calibration.
Figures & tables
Figure 2: Overview of TRACE and TRACE+. Given a prompt and a decoded answer, TRACE extracts token-level surprisal and predictive entropy from the decoding trace, localizes risk through early/global, position-decayed, local-peak, and shape/length operators, and converts the resulting risk into a monotone confidence score. TRACE+ further standardizes trace features and applies a lightweight logistic calibrator to produce a calibrated probability.
Method
MLQA
SVAMP
TriviaQA
TruthfulQA
Avg.
Brier ↓
AUROC ↑
Brier ↓
AUROC ↑
Brier ↓
AUROC ↑
Brier ↓
AUROC ↑
Brier ↓
AUROC ↑
Token- and sequence-level confidence baselines
First-Token Prob ( Chen et al., 2026 )
0.230
0.580
0.211
0.721
0.188
0.806
0.094
0.590
0.181
0.674
MeanProb ( Flores et al., 2025 )
0.228
0.692
0.196
0.782
0.206
0.774
0.094
0.615
0.181
0.716
Len-Norm LogP ( Bakman et al., 2024 )
0.226
0.694
0.192
0.778
0.198
0.783
0.094
0.619
0.178
0.719
SeqLogP / Total NLL ( Aichberger et al., 2026 )
0.174
0.812
0.146
0.791
0.182
0.811
0.093
0.619
0.149
0.758
Table 1: Main comparison on general generation calibration. All methods score the same decoded answer; beam-based methods additionally use beam statistics from the same prompt.
Model
Avg. Acc.
Best Non-TRACE
TRACE+
Improvement
Task Wins
Brier ↓
AUROC ↑
Brier ↓
AUROC ↑
Δ Brier ↑
Δ AUROC ↑
Brier
AUROC
Llama-3.1-8B-Instruct
0.396
0.141
0.747
0.102
0.819
+0.039
+0.072
3/3
3/3
Mistral-7B-Instruct-v0.3
0.438
0.160
0.763
0.146
0.790
+0.014
+0.027
3/3
2/3
Phi-3.5-MoE-Instruct
0.431
0.135
0.793
0.121
0.827
+0.014
+0.034
3/3
3/3
Qwen2-57B-A14B
0.371
0.169
0.714
0.143
0.819
+0.026
+0.105
3/3
2/3
Llama-3.1-70B-Instruct
0.592
0.104
0.819
0.098
0.844
+0.007
+0.024
3/3
2/3
Table 2: Cross-model generalization against the full non-TRACE baseline pool. Metrics are averaged over SVAMP, TriviaQA, and TQA-Gen. Best Non-TRACE reports the strongest value achieved by any non-TRACE estimator.
Figure 5: Local-risk aggregation analysis. Panel (a) shows real incorrect generations with localized decoding-risk peaks. Panel (b) compares mean aggregation with TRACE local-peak risk on the same examples.
Variant
Brier ↓
AUROC ↑
ECE ↓
TRACE (raw)
0.182
0.772
0.163
TRACE + scalar calib.
0.154
0.772
0.092
Learned (D,M,L)
0.140
0.778
0.054
TRACE+
0.137
0.792
0.054
Table 3: Bridge from TRACE to TRACE+, comparing calibration, learned operators, and richer trace features.
Variant
Main
Cross
Brier ↓
AUROC ↑
Brier ↓
AUROC ↑
TRACE+
0.137
0.792
0.120
0.817
+ seq. likelihood
0.139
0.787
0.121
0.816
w/o local entropy
0.138
0.786
0.121
0.812
w/o local surprisal
0.137
0.792
0.121
0.815
w/o trajectory slope
0.136
0.792
0.123
0.808
Table 4: Ablation of TRACE+ components in main and cross-model settings, reporting Brier and AUROC for feature removals and local-risk-only variants.
Method
Brier ↓
AUROC ↑
MARS ( Bakman et al., 2024 )
0.1730
0.7540
MARS + TRACE
0.1410
0.7750
TokenSAR ( Duan et al., 2024 )
0.1740
0.7380
TokenSAR + TRACE
0.1410
0.7810
TRACE+
0.1371
0.7920
TRACE+ + MARS
0.1376
0.7910
Table 5: Semantic fusion between TRACE and the semantic confidence estimators MARS and TokenSAR.
Target
Best Non-TRACE
TRACE+
Gain
Brier ↓
AUROC ↑
Brier ↓
AUROC ↑
Brier ↑
AUROC ↑
MLQA
0.264
0.813
0.347
0.814
-0.083
+0.001
SVAMP
0.144
0.843
0.140
0.842
+0.004
-0.001
TriviaQA
0.182
0.835
0.169
0.832
+0.013
-0.003
TruthfulQA
0.137
0.644
0.212
0.659
-0.075
+0.016
Table 6: Cross-task transfer against the strongest source-calibrated baseline on unseen held-out target tasks.
Appendix figures & tables15 assets
Supplementary material from the paper’s appendix.
Appendix
Mean risk
Spike gap
n
Mean risk
Spike gap
Incorr. rate
Low
Low
313
0.064
0.023
0.224
Low
High
312
0.090
0.182
0.554
Mid
Low
313
0.178
0.102
0.591
Mid
High
313
0.189
0.390
0.843
High
Low
313
0.377
0.139
0.850
High
High
312
0.321
0.454
0.904
Appendix
Table 7: Incorrect rates by mean risk and spike gap.
Feature set
AUROC ↑
Brier ↓
Mean risk
0.774±0.015
0.179±0.004
Max token risk
0.784±0.014
0.173±0.005
Spike gap
0.761±0.018
0.184±0.005
Mean risk + spike gap
0.808±0.013
0.164±0.005
TRACE local features
0.818±0.012
0.159±0.005
Appendix
Table 8: Predictive value of localized risk features.
Group
Features
Early entropy
First-3 entropy confidence
Global entropy
Mean entropy confidence
Length
Log answer length
Trajectory
Entropy slope
Local entropy
Decayed entropy confidence ( λ=2,4 )
Local surprisal
Decayed surprisal confidence ( λ=2,4 )
Appendix
Table 9: TRACE+ feature groups and definitions.
Method
Same
Single
Extra
Sem.
Calib.
Likelihood / entropy
✓
✓
✗
✗
✗
Beam-based scores
✓
△
✗
✗
✗
TokenSAR
✓
✓
✗
✓
✗
MARS
✓
✓
✗
✓
✗
SelfCheckGPT
✓
✗
✓
△
✗
Semantic Entropy
✗
✗
✓
✓
✗
Appendix
Table 10: Protocol comparison of confidence estimators by answer preservation, single-pass inference, extra generations, semantic modeling, and calibration.
Model
Task
Brier comparison
AUROC comparison
Best estimator
Base ↓
TRACE+ ↓
Δ↑
Best estimator
Base ↑
TRACE+ ↑
Δ↑
Llama-3.1-8B
SVAMP
High-Prob Token Rate
0.134
0.064
+0.071
Mean Token Entropy
0.939
0.974
+0.035
TriviaQA
First-Token Prob.
0.161
0.158
+0.003
First-Token Prob.
0.798
0.801
+0.003
TQA-Gen
First-Token Prob.
0.088
0.085
+0.003
First-Token Prob.
0.646
0.683
+0.037
Mistral-7B
SVAMP
SeqLogP / Total NLL
0.190
0.178
+0.012
TokenSAR
0.824
0.808
-0.016
TriviaQA
First-Token Prob.
0.180
0.165
+0.015
Max Token Entropy
0.799
0.823
+0.024
Appendix
Table 11: Task-level details for cross-model generalization.
Method
Req.
Brier ↓
AUROC ↑
MARS
–
0.173
0.754
TokenSAR
–
0.174
0.738
LARS
L
0.157
0.758
P(True)
E
0.178
0.715
Verbalized conf.
E
0.183
0.697
Internal-state probe
H+L
0.159
0.766
Appendix
Table 12: Comparison with uncertainty-estimation methods for answer-level confidence and ranking.
Model
Best Non-TRACE for Brier
Best Non-TRACE for AUROC
TRACE+
Estimator
Brier ↓
Δ Brier ↑
Estimator
AUROC ↑
Δ AUROC ↑
Brier ↓
AUROC ↑
Llama-3.1-8B-Instruct
High-Prob Token Rate
0.141
+0.039
Mean Token Entropy
0.747
+0.072
0.102
0.819
Mistral-7B-Instruct-v0.3
SeqLogP / Total NLL
0.160
+0.014
TokenSAR
0.763
+0.027
0.146
0.790
Phi-3.5-MoE-Instruct
SeqLogP / Total NLL
0.135
+0.014
SeqLogP / Total NLL
0.793
+0.034
0.121
0.827
Qwen2-57B-A14B
SeqLogP / Total NLL
0.169
+0.026
Max Token Entropy
0.714
+0.105
0.143
0.819
Llama-3.1-70B-Instruct
SeqLogP / Total NLL
0.104
+0.007
SeqLogP / Total NLL
0.819
+0.024
0.098
0.844
Appendix
Table 13: Best averaged non-TRACE estimators in the cross-model study.
Method
AUROC ↑
Localized errors
Global errors
Loc.
Glob.
TPR@5% ↑
TPR@10% ↑
TPR@20% ↑
pAUC@10% ↑
TPR@10% ↑
SeqLogP
0.763
0.767
0.190
0.395
0.555
0.572
0.330
WindowEnt
0.778
0.754
0.245
0.378
0.664
0.599
0.371
MARS
0.778
0.789
0.308
0.494
0.637
0.620
0.330
TRACE
0.796
0.796
0.346
0.471
0.691
0.618
0.474
TRACE+
0.805
0.744
0.272
0.521
0.692
0.648
0.423
Appendix
Table 14: Localized- and global-error discrimination.
Variant
Avg. AUROC ↑
Position-decayed entropy D
0.751
Local-window entropy M
0.754
Length-normalized surprisal L
0.758
TRACE (D+M+L)
0.772
Learned (D,M,L)
0.776
Appendix
Table 15: Operator-level analysis of TRACE components and their fusion.
Param.
Values tested
Brier ↓
AUROC ↑
α
0.2 – 0.6
0.154 – 0.155
0.767 – 0.775
β
0.2 – 0.6
0.154 – 0.156
0.767 – 0.774
γ
0.1 – 0.4
0.153 – 0.156
0.769 – 0.772
λ
1,2,4,8
0.154 – 0.156
0.763 – 0.774
w
2,4,6,8
0.153 – 0.156
0.769 – 0.774
ρ
0,.125,.25,.5,.75,1
0.152 – 0.160
0.766 – 0.773
Appendix
Table 16: Hyperparameter sensitivity analysis.
Task
Original ↑
Position-shifted ↑
MLQA
0.786
0.769
SVAMP
0.825
0.809
TriviaQA
0.835
0.811
TruthfulQA
0.641
0.630
Average
0.772
0.755
Appendix
Table 17: AUROC under risk-position perturbation.
Method
Number/ Arithmetic
Entity/ Span
Factual Claim
Overall Localized
(n=92)
(n=438)
(n=65)
(n=595)
SeqLogP
0.845
0.801
0.528
0.757
WindowEnt
0.868
0.801
0.556
0.770
MARS
0.879
0.796
0.589
0.768
TRACE
0.871
0.821
0.569
0.784
TRACE+
0.927
0.836
0.607
0.805
Appendix
Table 18: AUROC on localized-error subtypes.
Calib.
Metric
SeqLogP
WinEnt
MARS
TRACE
TRACE+
5%
Brier
0.154
0.182
0.188
0.178
0.156
AUROC
0.758
0.754
0.754
0.772
0.739
10%
Brier
0.152
0.175
0.183
0.171
0.148
AUROC
0.759
0.756
0.755
0.773
0.761
20%
Brier
0.148
0.166
0.177
0.161
0.140
AUROC
0.762
0.759
0.757
0.775
0.778
Appendix
Table 19: Calibration-size robustness in Brier/AUROC.
Method
ECE ↓
Brier ↓
AUROC ↑
R@10 ↓
R@50 ↓
R@90 ↓
SeqLogP
0.036
0.146
0.826
0.262
0.352
0.474
WinEnt
0.058
0.149
0.826
0.256
0.357
0.476
MARS
0.048
0.159
0.783
0.242
0.357
0.479
TRACE
0.059
0.144
0.830
0.251
0.350
0.473
TRACE+
0.013
0.134
0.877
0.239
0.331
0.474
Appendix
Table 20: Reliability and selective-prediction results.
Method
Short
Medium
Long
SeqLogP
0.135/0.733
0.151/0.761
0.142/0.717
WinEnt
0.144/0.737
0.161/0.748
0.154/0.740
MARS
0.157/0.710
0.172/0.781
0.172/0.752
TRACE
0.140/0.750
0.155/0.781
0.151/0.741
TRACE+
0.128/0.761
0.137/0.792
0.136/0.758
Appendix
Table 21: Length-stratified robustness in Brier/AUROC.
Large language models (LLMs) can generate fluent reasoning traces that nevertheless lead to incorrect answers, making response-level uncertainty estimation important for abstention, human review, and adaptive compute allocation. Existing approaches generally fall into three categories: passive single-trace methods use token-level confidence signals, sampling-based methods compare multiple complete traces at higher generation cost, and active prefix-based methods probe partial traces to study answer stabilization or preference transitions. However, none actively re-elicits an answer from a completed reasoning trace to measure its consistency with and support for the original answer. To address this gap, we introduce Trace-Conditioned Answer Consistency (TrAC), a correctness-supervised uncertainty quantification framework that combines active and passive signals anchored to one completed reasoning trace. Its active component, Prefix-Conditioned Elicitation (PCE), re-elicits a short answer conditioned on the completed trace and represents both its consistency with the original answer and its token-level probabilistic support. Its passive component, Trace Uncertainty Profile (TUP), summarizes how token-level uncertainty evolves throughout the original generation without additional decoding. A lightweight head then integrates the two representations into a response-correctness score. Across five mathematical reasoning benchmarks and three LLM families, TrAC improves macro AUROC by 1.8% and reduces AURC by 3.4% relative to eight-sample self-consistency, while using one complete reasoning trace and a short cached answer probe. When eight samples are already available, augmenting sample consensus with re-elicitation further improves macro AUROC by 4.3% and reduces AURC by 8.3%, without additional full-trace generation.
Reasoning language models can solve increasingly complex tasks, but struggle to produce the calibrated confidence estimates necessary for reliable deployment. Existing calibration methods usually depend on labels or repeated sampling at inference time, making them impractical in many settings. We introduce a method for unsupervised confidence calibration of reasoning LLMs when only a single generation is available at inference time. Our approach uses offline sampling on unlabeled data to derive a self-consistency-based proxy target, then distills this signal into a lightweight deployment-time confidence predictor. In a broad evaluation across 5 math and question-answering tasks using 9 reasoning models, our method substantially outperforms baselines, including under distribution shift, and improves downstream performance in selective prediction and simulated downstream decision-making.
The maximum softmax probability (MSP) represents a default approach when evaluating uncertainty quantification for language model generation with structured output. Although cheap, it is often miscalibrated. Methods that probe the model's internal activations feed raw hidden states into opaque classifiers, reading activations as static snapshots and leaving implicit the layer-wise trajectory by which a representation is formed. Yet, similar endpoints can arise from very different paths, and how evidence accumulates, reinforces, or reverses across depth might reveal uncertainty that final probabilities obscure. We extract eleven scale-invariant geometric features, tracing the cumulative path of per-layer MLP updates, and feed them to a sparse linear probe. The probe outperforms MSP under selective abstention, with gains scaling with baseline miscalibration up to 21 AURC points. Because every feature has a closed-form geometric meaning, the probe's coefficients trace how and where along depth errors take shape -- which layers commit prematurely, which contradict the running state, where trajectories drift away from their endpoint.