Certo is a small non-generative decision model (Qwen3-4B): it scores candidate actions from their text and returns a probability, instead of generating an answer. The accurate design reads the state, the rules, and each candidate together (a joint scorer), so cost grows with the menu. Independent encoding lets each candidate be encoded once and reused across states (about 5x cheaper at 77 candidates), but separates state from candidate. We ask how much rule-sensitivity survives that move, and whether it can be trained back. Four experiments on Certo: (1) the tested conversion to cacheable scoring loses rule-sensitivity (recall@1 1.00 -> 0.24) while the joint scorer holds 1.00, and a shortlist+rerank rescue fails; (2) targeted counterfactual supervision restores strong performance on held-out synthetic rule tasks (paraphrase, counterfactual, composition; reproducible across seeds), though we do not isolate whether predictions depend on the supplied rule; (3) on real rules the added benefit is not established -- after fixing a truncation confound, the joint scorer wins significantly on the short tier (0.861 vs 0.500) and directionally on the hard tier (0.655 vs 0.483, n=29); (4) a matched cross-domain real-prose mixture did not help and reduced contract accuracy (-9.3, -16.2 points). A cacheable encoder can be made rule-sensitive on its training distribution, but transfer to unseen-source real rules is not established; the joint scorer keeps an edge at the cost of caching.
Figures & tables
routing
rule-sensitive
cand. emb.
scorer
K5
K20
K77
K5
K17
latency
reusable
Apair (joint)
0.938
0.853
0.652
1.00
1.00
52–242 ms
no
A10 (joint head)
0.953
—
—
1.00
—
47 ms ( K≤10 )
no
B (single-vector)
0.802
0.573
0.422
0.542
0.240
46 ms
yes
C (multi-vector)
0.242
0.048
0.013
0.295
0.080
46 ms
yes
Table 1: Experiment 1: recall@1 by scorer, with p50 latency. Reusable : candidate embeddings reusable across states; the joint scorer can still reuse the state prefix within one state via prefix caching. The rule-sensitive slice has 17 candidates, so its second column is K17 (the full menu), not K20.
block
B (before)
control
treatment
treat. − control (95% CI)
unseen wording (exc. / thr.), n=300
0.51 / 0.55
0.27 / 0.59
0.99 / 1.00
+0.563[0.507,0.620]
counterfactual, 150 pairs
0.55
0.56
0.99
+0.430[0.377,0.483]
new composition, n=300
0.42
0.48
0.73
+0.250[0.187,0.313]
novel type (multicond. / prec.), n=300
0.43 / 0.64
0.61 / 0.60
1.00 / 0.73
+0.263[0.217,0.313]
counterfactual both-correct
0.30
0.29
0.987
—
routing (retention)
0.430
0.453
0.420
−0.033[−0.077,+0.013]
Table 2: Experiment 2: held-out generalization suite, recall@1. Slashes give the two families within a block; the CI is for the pooled block. Bootstrap grouping: by pair for the counterfactual block ( 300 items, 150 pairs), by item otherwise ( n=300 ). Every block gain over both baselines excludes zero.
slice
n
B
control
treatment
Apair (joint)
short tier, choice (fits)
36
0.500
0.472
0.500
0.861
hard tier, choice (fits)
29
0.414
0.379
0.483
0.655
hard tier, choice (truncated)
38
0.211
0.132
0.105
0.079
short tier, yes/no (fits)
24
0.958
0.958
0.917
0.958
Table 3: Experiment 3: real rule-sensitive choice, recall@1, rescored within the trained window. On truncated items the joint scorer falls furthest, exactly as the candidate-dropping mechanism predicts.
slice
n
rule-aware start
control (synthetic)
treatment (+real prose)
contracts (fit ≤ 640)
332
0.870
0.870
0.777
contract-qa (all fit)
80
0.988
0.963
0.800
contracts (truncated)
144
0.569
0.562
0.556
synthetic rules (retention)
1200
0.897
0.928
0.903
routing (retention)
400
0.417
0.380
0.388
Table 4: Experiment 4: real-prose pilot, recall@1. Replacing part of the synthetic mixture with privacy-policy prose reduced unseen-source contract accuracy ( −9.3 and −16.2 points); retention on synthetic rules and routing is approximately unchanged (we do not run an equivalence test).