Organizations: Department of Mechanical Engineering, The University of Tokyo, Tokyo, Japan · School of Mechanical Engineering and Automation, Harbin Institute of Technology, Shenzhen, China · Faculty of Mechanical Engineering, RWTH Aachen University, Aachen, Germany · Department of Computer Science and Engineering, HKUST, Hong Kong SAR, China
Learning systems adapt quickly inside a familiar family of models. The harder step comes earlier: deciding, from observations that could be noise, an exception, a change within the family or structure outside it, whether opening a richer family is worth its cost. We treat this as a costly sequential decision: prediction failure must be turned into structural evidence, evidence into a value of expansion, and value into action. The Structural Revision Environment produces matched failures from each source, varies the price of expansion and the remaining horizon independently of the evidence, and admits exact Bayesian calculations and an exact normative solution of the one-shot decision. Its solution shows that revision is a value boundary and not an evidence threshold: one history has different optimal actions under different prices, horizons and announced queries, the boundary between local repair and expansion is set by the inputs a rule predicts and a repair cannot cover, and belief in the richer family crosses long before the decision does. Transformers trained in the environment reproduce this boundary from utility alone. Language models of three post-training lineages carry a failure-sensitive signal in their predictions that is not reflected in their revision decisions, and given the gain of expanding they read it without weighing it against price and horizon. Three models allowed to reason weigh the stated gain in the reference's proportions and still do not turn the history into an estimate of what expansion would buy. Controlled post-training of the meta-trained learners moves the prior and the sharpness of predictions, and neither moves the criterion.
Figures & tables
Figure 1: The Structural Revision Environment and its normative solution. Left: four episodes, one per class, that share one query schedule and one set of label flips; framed squares mark the observations that the current family fails to predict. S is the exception set, b the threshold, τ the change point and u the upper end of the interval. Middle: the nested hypothesis spaces, the four actions with their prices, and the optimal action for one history with two anomalous inputs at cE=4 and three horizons. Right: the expansion advantage ΔQ∗ as a surface over the reference’s belief in the richer family (its posterior log-odds) and amortized price ( H=20 , cp=0.5 ), above the regions of the optimal action; the dark line is ΔQ∗=0 , and the three dots are the three horizons of the middle panel at their amortized prices.
quantity
definition
when it grows
Shannon surprise St
−logqt(yt∣xt)
Whenever an improbable label arrives, including under a correct family
Structural evidence Et
∑s≤tlogU(ys∣xs,z<s)−suph∈F0logph(z≤t)
When the residuals of F0 are compressible
Value of expansion ΔQ∗(st)
Q∗(st,expand)−maxa=expandQ∗(st,a)
Table 2
class
mechanism
what it isolates
noise
A threshold rule. Failures come from label flips alone
Failures that do not reproduce. A repeated query returns the expected label, and any revision is wasted
exception
A threshold rule with one or two inputs whose label is inverted
Failures that reproduce and do not generalize. One slot per input suffices, and expansion buys nothing more
drift
The threshold moves once during the episode
Table 3
Figure 2: Learners trained in the environment weigh the price of expanding as the reference does; language models value a stated gain only when they reason. (a) The meta-trained Transformers (five regimes, three seeds each) and the bounded reference lie on the reference (zoom); at the gain level, eleven language models lie in the fixed-threshold corner with the e-process trigger, and the two 7B OLMo 3 Think models respond to the gain, the price and the rounds left by a few hundredths each, too little for their proportion to be identified (Appendix Q ); with reasoning (Qwen 3 8B at its own close, OLMo 3 Think and Think SFT at 8,192 tokens) the three models move onto the reference’s direction while their slopes grow by an order of magnitude; at the history level, twelve of thirteen models do not respond to the benefit. (b) At rule log-odds 0 to 2, the Transformer trained on utility alone (mean over three seeds, band the range) expands less as the amortized price rises, as the reference does, and the fixed trigger, OLMo 3 Instruct and OLMo 3 Think are nearly flat. (c) At the decision points where the reference expands, the share at which each language model’s label predictions have moved toward the rule, and the share at which its decision is expand, with 95% intervals.
Figure 3: What responds to a failure. (A) Matched-pair κ , the share of the Bayes shift toward the interval rule that a learner’s predictions make on a rule episode relative to its matched threshold episode, against the number of distinct anomalous inputs. Thin lines: Bayes predictors indexed by their prior on the rule class. The three meta-trained Transformers lie near the members with their training base rates. Bands: the 14 language models of three lineages, each model the mean over three renderings, each band the range within a lineage, with eight threshold-episode demonstrations in the context. (B) The shift in log-odds between a rule episode and its matched threshold episode at inputs placed by their distance from the end of the rule’s region, on 110 matched pairs. Right: the step at the boundary, the coefficient of 1[x>u] in a regression of the shift on u−2≤x≤u+3 , with 95% bootstrap intervals over episodes, one row per model.
Appendix figures & tables21 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 4: The value of waiting. (A) The one-shot grid of Figure 1 ( H=20 , cp=0.5 , informative announcement): cells where the announced label has value are shaded by its VOI, the others by the action committed to. At any co the reference defers exactly where VOI exceeds co (F1). (B) Mean VOI over the grid against the number k of announced labels revealed before committing, for informative and uninformative extra queries (F2). (C) The reference’s mean and Kaplan–Meier median expansion step in rule episodes and the bound ( 13 ), against the rate ϱ of discriminating queries at repeat probabilities r=0.1 , 0.35 and 0.6 (F3). Bars: one bootstrap standard error over episodes.
Figure 5: The Qwen 3 pair against the family indexed by the prior. Matched-pair κ of Qwen 3 8B Base and Qwen 3 8B against the evidence level m , with the test episode alone (A), after eight threshold-episode demonstrations (B) and after eight demonstrations from the generative mixture (C), next to E1 under rule priors from 0 to 0.25 (grey, labelled in C). Lines are medians over the null-equivalent pairs averaged over the three renderings, and bands the range over the renderings.
Table 1: Meta-decision regret against the reference at the default cell. T=40 , cE=4 , cp=0.5 (a held-out (cE,T) cell of the small Transformers’ training split); nats per episode, prior-weighted over the four classes (noise 0.40, exception 0.20, drift 0.15, rule 0.25); regret =U(reference)−U(policy) on the same 1 200 episodes (300 per class). Small Transformers: mean over three seeds; their SE is the per-run paired SE averaged over the seeds. Path labels: reference actions on the states beyond the expansion. Latency: Kaplan–Meier median of the expansion step and the censored share. Patch only: the probability of patching without expanding in exception episodes. Grid mean: regret averaged over the 120 price cells.
Table 2: Meta-decision regret against the reference at the default cell: rows and columns not shown in Table 1 . Top: the baselines left out of Table 1 , with its columns; the grid mean is given with the threshold fixed and with the threshold re-tuned in every cell. Bottom: the remaining columns, for every row.
Figure 6: Post-training separates the prior from predictive sharpness. (a) Contrasts from the 2×2 post-training intervention after 20,000 steps. Points are means over three learners; bars are 95% bootstrap intervals, with filled markers excluding zero. κ is rank-mapped to remove sharpening. (b) Recovery of the implied prior and rule-log-odds decodability after restoring the original data and objective.
Figure 7: Adapter branches on two base models. Paired contrasts of the four branches on OLMo 3 7B Base (circles) and Llama 3.1 8B Base (squares) without demonstrations, in the layout of Figure 6 (a), with the probe readouts replaced by surprise leakage and by the step at the boundary divided by the sharpness. Bars are 95% intervals of one episode bootstrap shared by all models.
Figure 8: Detection and commitment. (A) Shift of the belief crossing mB∗ , exact and in closed form, and of the decision crossing mD∗ (median and 10th to 90th percentile over prices, one or two observations per anomalous input) against the size of the rule library, relative to ∣R∣=120 (J1). (B) Search delay, the reference’s expansion step, and commitment delay, the step at which its posterior puts more than half of the mass of the rule class on the true rule, in 300 rule episodes whose rules belong to every library: means with one bootstrap standard error, and Kaplan–Meier medians (J2).
Table 3: Architectures at the generative base rate: mean ( ± s.d.) over three seeds. The Bayes predictor is E1 , the Bayes predictor of the training prior of all three learners (intervals: 95% over episodes).
Figure 9: Recovery after the lift. (A) Excess log-loss over E1 of the Transformer, the GRU and the selective state-space learner at the k -th prediction after each learner’s own lift on rule episodes, k=0 being the prediction of the anomalous observation that triggers it (mean over three seeds; bands span the seeds’ 95% intervals over episodes). (B) The log-loss itself, against E1 aligned at its own lift.
Figure 10: Steering along the probe and encoding directions. Dashed lines mark k=±1 , where the pre-specified criterion is read; bands are 95% intervals over models, then episodes. From left to right: lift-time shift on the base learners; their in-family controls, where negative values are false lifts; lift latency of the post-trained branches; the branches’ in-family controls.
Table 4: The generative process stated in the prompt (R1S) against R1 and the length control R1L, at the history level in order A: the probability of expanding on the noise histories of the matched groups, from which the other classes differ by at most 0.05, its gap between rule histories and their boundary-matched threshold histories, and the benefit coefficient βb on the same-history points, with 95% intervals from the paired cluster bootstrap. No cell meets the criterion, in either action order.
null-equivalent pairs
moved toward the rule
expand most probable
model
median κ
expand, D−D′
decision prefix
implicit
all
moved
not moved
label mass
OLMo 3
OLMo 3 stage 1
0.02 [ 0.01 , 0.03 ]
0.000
0.61 [ 0.53 , 0.69 ]
0.65
0.56 [ 0.49 , 0.63 ]
0.55
0.58
0.88
OLMo 3 Base
0.08 [ 0.06 , 0.12 ]
−0.001
0.62 [ 0.54 , 0.69 ]
0.69
0.00 [ 0.00 , 0.00 ]
0.00
0.00
0.92
OLMo 3 Instruct SFT
0.25 [ 0.22 , 0.29 ]
−0.004
0.54 [ 0.46 , 0.62 ]
0.65
0.00 [ 0.00 , 0.00 ]
0.00
0.00
0.98
OLMo 3 Instruct DPO
0.40 [ 0.33 , 0.47 ]
−0.006
0.53 [ 0.45 , 0.60 ]
0.64
0.00 [ 0.00 , 0.00 ]
0.00
0.00
1.00
Appendix
Table 5: The prediction read on the decision-context prefix, under R1 in order A: on the null-equivalent pairs the median matched-pair κ and the difference in the probability of expanding between D and D′ , and at the 1,603 points of Figure 2 (c) the share at which the prediction has moved toward the rule, on the decision-context prefix and in the implicit protocol, the share at which expand is the most probable action, at all points and where the prediction has or has not moved, and the probability mass on the two label symbols. The intervals are 95%, over pairs for κ and from the cluster bootstrap elsewhere.
Table 6: Training runs of the small learners. Learning rates rise linearly over the warm-up and then follow a cosine to 10% of the peak, except in post-training, where they stay constant. Every run draws fresh episodes at each step except the three BC regimes, which read the label set.
model
stage
repository
OLMo 3, 7 billion parameters
OLMo 3 stage 1
pre-training, stage 1
allenai/Olmo-3-1025-7B
OLMo 3 Base
base
allenai/Olmo-3-1025-7B
OLMo 3 Instruct SFT
SFT
allenai/Olmo-3-7B-Instruct-SFT
OLMo 3 Instruct DPO
DPO
allenai/Olmo-3-7B-Instruct-DPO
OLMo 3 Instruct
RLVR
allenai/Olmo-3-7B-Instruct
Appendix
Table 7: Sources of the language models: each repository name links to the revision that was read.
matched-pair κ
model
m=1
m=3
m=5
leakage
boundary step
far outside
false lift
s0
Base
OLMo 3 32B Base
0.30
0.29
0.44
0.137
0.25 [ −0.51 , 1.07 ]
2.32
0.28
0.614
OLMo 3 Base (7B)
0.41
0.40
0.58
0.110
0.24 [ −0.52 , 1.08 ]
2.25
0.26
0.717
Think SFT
OLMo 3 32B Think SFT
0.28
0.33
0.49
0.149
0.22 [ −0.63 , 1.10 ]
2.28
0.32
0.629
Appendix
Table 8: The 32B end points against their 7B counterparts in the implicit protocol, with the test episode alone. Matched-pair κ at m=1 , 3 and 5 anomalous inputs; surprise leakage; the step at the boundary in log-odds with its 95% interval over episodes; the response far outside the region; the false lift on the matched threshold episode; the calibration slope s0 . The 7B counterpart of OLMo 3.1 32B Think is OLMo 3 Think. The Bayes predictor E1 has κ=1 at every m , leakage 0, a step of 1.12 [ 0.88 , 1.37 ], a far-outside response of −0.12 and a false lift of 0.02.
Table 9: The 32B end points against their 7B counterparts in the explicit protocol (chat format with the path read-out, rendering R1, order A, natural states; 95% intervals from the paired cluster bootstrap). History level: the benefit coefficient βb , the spread of the probability of expanding over the five classes, its gap between rule histories and their matched threshold histories, and the change in the probability of deferring under a discriminating announcement; the failure-level column gives βb at that level. Gain level: price and horizon sensitivity, n.i. where not identified. The verdicts are those of Section 5.2 at the history level (constant, or partial when a small benefit response holds in some cells) and at the gain level (read alone, weighed in the reference’s proportion, mixed, or n.i.). OLMo 3 Think DPO was not read at 7B.
Table 10: Qwen 3 8B under a thinking budget: the gap in the probability of expanding between rule histories and their matched threshold histories at the history level, and price and horizon sensitivity at the gain level, by budget and action order. Thinking off is the budget 0. Every interval of the gap lies within [−0.006,0.007] ; the sensitivities carry 95% intervals from the cluster bootstrap.
thinking tokens at the close
level, action order
chains
share closed
median
90th percentile
Matched groups
history level, order A
600
0.930 [ 0.913 , 0.947 ]
14,748
21,848
history level, order B
600
0.963 [ 0.947 , 0.979 ]
13,452
19,993
Same-history design
history level, order A
366
0.970 [ 0.952 , 0.985 ]
14,124
20,656
Appendix
Table 11: Closing of the thinking of Qwen 3 8B in the natural-closing arm, by set. Share closed within the cap of 24,576 tokens, with its 95% interval, and the median and 90th percentile of the number of thinking tokens at the close.
Table 12: Qwen 3 8B at its own close, history level, matched groups pooled over the two action orders (closed chains): the probability of each action by class, the number of chains whose first choice is expand, and the share of points at which the reference’s best action is expand. The probability of expanding carries its 95% interval.
Table 13: The dose of reasoning for the three models. History level: the gap in the probability of expanding between rule histories and their boundary-matched counterfactuals (matched groups, one chain per point), in each action order. Gain level: price and horizon sensitivity on the 1,098 same-history points, order A; n.i.: not identified. Qwen 3 8B is read at prefixes of the chains of its natural-closing arm and at its own close, the OLMo 3 Think models at prefixes of their chains up to their cap of 8,192 tokens. 95% intervals from the cluster bootstrap; the thinking-off order-B history-level cell of OLMo 3 Think fails the validity condition.
Figure 11: Reasoning recovers the valuation of a stated gain, but not structural evidence from the history. (a) Price and (b) horizon sensitivity at the gain level against the thinking budget, with 95% intervals and the reference’s proportion, 0.7 to 1.3, as a band. (c) The history-level gap in the probability of expanding between rule histories and their boundary-matched counterfactuals, in orders A (filled) and B (hollow), against the criterion of 0.10; grey points are not decisions, and natural is the own close of Qwen 3 8B.
Recent benchmarks reveal that despite strong reasoning capabilities, large language models (LLMs) still struggle to faithfully apply complex contextual knowledge. These failures are often not wholesale reasoning collapses: in context-rich tasks, models may follow the central reasoning path while missing peripheral, persistent, or format-sensitive requirements.
Bayesian accounts of in-context learning face a direct objection: exact posterior predictives for exchangeable data are invariant to task-preserving order, yet transformers change next-token probabilities when the same examples are serialized differently. We show this objection targets a structural invariant rather than the quantity scoring online prediction. For any Bayesian reference, excess prequential code length is exactly cumulative predictive KL. For unordered support sets that must be serialized, the expected regret of a single admissible ordering decomposes into that of the order-averaged predictor plus an order-averaging gain. Exchangeability violations are therefore not binary refutations; they are priced by log loss. We instantiate the theory with KT/Dirichlet finite-alphabet prediction and coarsened Bayesian linear-regression (BLR) predictive distributions. On Qwen2.5-7B/14B, floored candidate distributions at support 256 have one-step excess code lengths of 0.020/0.011 bits for Bernoulli and 0.039/0.022 bits for four-way categorical prediction, with candidate mass above 0.999; coarsened BLR continuations increasingly match the posterior-predictive digit distribution as support grows. A frequentist plug-in baseline sharpens the reading: the predictive distributions sit closer to the Bayesian posterior predictive than to the maximum-likelihood plug-in, by a margin largest at small support, where the plug-in is degenerate, and vanishing as the references converge. Position interventions and a from-scratch ablation localize order sensitivity to the positional encoding, activation patching tests causal use of decoded sufficient statistics, and permutation mixtures quantify the downstream log-loss cost of arbitrary orderings. Transformers need not realize exchangeable posterior predictives for every serialization to be Bayes-competitive prequential predictors.
Reinforcement learning (RL) has become a standard tool for post-training language models on reasoning tasks, where the policy is updated by reward feedback while exploring the space of responses. Despite its empirical success, theoretical understanding of RL post-training remains limited, in particular of why on-policy exploration combined with a neural reward model is effective. In this paper, we address this question by modeling the reward as a hierarchical function on the response space: the reward consists of infinitely many local components, each of which becomes relevant only after the preceding ones have been resolved. We show that a natural Transformer-based actor--critic algorithm, which alternates between sampling from the current KL-regularized policy, fitting a Transformer critic to the observed rewards, and updating the policy, achieves the minimax optimal rates in the query budget and in the regularization strength up to logarithmic factors, and is minimax optimal for a fixed number of prompts. In contrast, we prove that sampling from the fixed reference distribution, as in offline reward modeling, can limit regret decay to a logarithmic rate. These results show that on-policy exploration progressively zooms in on the region where the reward is concentrated, and quantify its benefit for RL post-training.