Certo is a small non-generative decision model (Qwen3-4B): it scores candidate actions from their text and returns a probability, instead of generating an answer. The accurate design reads the state, the rules, and each candidate together (a joint scorer), so cost grows with the menu. Independent encoding lets each candidate be encoded once and reused across states (about 5x cheaper at 77 candidates), but separates state from candidate. We ask how much rule-sensitivity survives that move, and whether it can be trained back. Four experiments on Certo: (1) the tested conversion to cacheable scoring loses rule-sensitivity (recall@1 1.00 -> 0.24) while the joint scorer holds 1.00, and a shortlist+rerank rescue fails; (2) targeted counterfactual supervision restores strong performance on held-out synthetic rule tasks (paraphrase, counterfactual, composition; reproducible across seeds), though we do not isolate whether predictions depend on the supplied rule; (3) on real rules the added benefit is not established -- after fixing a truncation confound, the joint scorer wins significantly on the short tier (0.861 vs 0.500) and directionally on the hard tier (0.655 vs 0.483, n=29); (4) a matched cross-domain real-prose mixture did not help and reduced contract accuracy (-9.3, -16.2 points). A cacheable encoder can be made rule-sensitive on its training distribution, but transfer to unseen-source real rules is not established; the joint scorer keeps an edge at the cost of caching.
Figures & tables
routing
rule-sensitive
cand. emb.
scorer
K5
K20
K77
K5
K17
latency
reusable
Apair (joint)
0.938
0.853
0.652
1.00
1.00
52–242 ms
no
A10 (joint head)
0.953
—
—
1.00
—
47 ms ( K≤10 )
no
B (single-vector)
0.802
0.573
0.422
0.542
0.240
46 ms
yes
C (multi-vector)
0.242
0.048
0.013
0.295
0.080
46 ms
yes
Table 1: Experiment 1: recall@1 by scorer, with p50 latency. Reusable : candidate embeddings reusable across states; the joint scorer can still reuse the state prefix within one state via prefix caching. The rule-sensitive slice has 17 candidates, so its second column is K17 (the full menu), not K20.
block
B (before)
control
treatment
treat. − control (95% CI)
unseen wording (exc. / thr.), n=300
0.51 / 0.55
0.27 / 0.59
0.99 / 1.00
+0.563[0.507,0.620]
counterfactual, 150 pairs
0.55
0.56
0.99
+0.430[0.377,0.483]
new composition, n=300
0.42
0.48
0.73
+0.250[0.187,0.313]
novel type (multicond. / prec.), n=300
0.43 / 0.64
0.61 / 0.60
1.00 / 0.73
+0.263[0.217,0.313]
counterfactual both-correct
0.30
0.29
0.987
—
routing (retention)
0.430
0.453
0.420
−0.033[−0.077,+0.013]
Table 2: Experiment 2: held-out generalization suite, recall@1. Slashes give the two families within a block; the CI is for the pooled block. Bootstrap grouping: by pair for the counterfactual block ( 300 items, 150 pairs), by item otherwise ( n=300 ). Every block gain over both baselines excludes zero.
slice
n
B
control
treatment
Apair (joint)
short tier, choice (fits)
36
0.500
0.472
0.500
0.861
hard tier, choice (fits)
29
0.414
0.379
0.483
0.655
hard tier, choice (truncated)
38
0.211
0.132
0.105
0.079
short tier, yes/no (fits)
24
0.958
0.958
0.917
0.958
Table 3: Experiment 3: real rule-sensitive choice, recall@1, rescored within the trained window. On truncated items the joint scorer falls furthest, exactly as the candidate-dropping mechanism predicts.
slice
n
rule-aware start
control (synthetic)
treatment (+real prose)
contracts (fit ≤ 640)
332
0.870
0.870
0.777
contract-qa (all fit)
80
0.988
0.963
0.800
contracts (truncated)
144
0.569
0.562
0.556
synthetic rules (retention)
1200
0.897
0.928
0.903
routing (retention)
400
0.417
0.380
0.388
Table 4: Experiment 4: real-prose pilot, recall@1. Replacing part of the synthetic mixture with privacy-policy prose reduced unseen-source contract accuracy ( −9.3 and −16.2 points); retention on synthetic rules and routing is approximately unchanged (we do not run an equivalence test).
When a language model follows an in-context conditional rule such as "if P(x) then A else B," does it assemble a runtime circuit with one module that tests the predicate and another that routes the answer? We probe this with activation patching under a four-donor design whose two swapped-rule donors make the condition and the answer word disagree, so each layer reveals which of the two it carries. Across three open models from two families and six languages sharing one fixed item bank, a mid-stack residual band carries the predicate's truth value: patching it reroutes the answer with predicate-outcome flip near 1.0 and mapping flip near 0.0, meeting a strict pre-specified isolation criterion in 17 of 18 cells, and the same localization holds across five predicate families. The router shows the opposite profile. A learned subspace flips A and B near-perfectly within the trained pair yet transfers to a new pair at approximately 0 in every model, while in Gemma-3-4B (the only model probed cross-lingually) it transfers at approximately 0.98 to the same pair in other languages. Under every probe we ran, the router direction is token-bound and non-transferable (largely answer-readout in Gemma, pair-specific in Qwen) rather than an abstract routing module. Test is modular; under these probes, route is not.
Luxshan Thavarasa, Sivasuthan Sukumar
Independent Researcher, Colombo, Sri Lanka · Department of Computer Science and Engineering, University of Moratuwa, Sri Lanka
Large language models are highly capable of answering difficult questions by retrieving, recombining, and attending to information in long contexts. For agentic tasks, an additional capability is required: the preservation of an exact state while repeatedly applying rules. We find that this reliability is absent across language models. To demonstrate, we query 126 leading model variants with the task of counting a long string of repeated characters, and we find they all cannot accurately count above a model-dependent, syntax-sensitive counting capacity threshold. Failures are abrupt and persist even with increasing model size, inference time computation, and external tool. Mechanistic probing indicates that models use a finite number of internal states to mimic counting as a rule and fail once these states are exhausted. Furthermore, such states are the basis for performing complex tasks beyond counting. These results indicate that fundamentally new model architectures are required for autonomous agents to achieve truly reliable rule following capabilities.
Tianxiang Dai, Jonathan Fan
Department of Electrical Engineering, Stanford University, Stanford, CA 94305.
Dedicated decision models such as Jev map unstructured language to probability distributions over finite choices, allowing their outputs to directly route requests, select tools, and trigger actions. Yet real-world inputs rarely arrive in isolation: they come with background details and surrounding context. We find that short additions that fit naturally into this context can nevertheless redirect an otherwise correct decision, even when the correct answer remains unchanged. To study this behavior, we fix a wrong target option for each initially correct item and use the model's option probabilities to refine fluent context additions while preserving the source, question, choices, and gold answer. Within 64 accepted target evaluations, the optimizer identifies contexts that redirect Jev on 312 of 508 initially correct decisions (61.4%); in 229 cases, Jev assigns at least 0.7 probability to the fixed wrong option. Across seven datasets, three additional decision systems show targeted flip rates of 64.9%-73.2% on decisions they initially answer correctly. Taken together, these results expose a pronounced fragility in current decision models: short, ordinary-looking context can shift a correct choice to a high-confidence wrong one. Because these models turn language directly into downstream choices, this sensitivity raises concerns about treating their probability outputs as reliable decision interfaces.