SanSi: A Looped Typed Decision Model for System 1.5 Thinking
Organizations: University of Minnesota
Abstract
Typed decision models answer a declared question without generating text: a decision head returns a probability for each of the declared options in a single forward pass. A single pass is fast, intuitive System 1 thinking. We study what lies between one pass and generated reasoning: looping, in which the same layers are recursively applied several times before one typed readout. Each loop lets the model revise its hidden state before it commits to an answer, without generating a token; we call this System 1.5 thinking. We propose SanSi, which turns a pre-trained looped language model into a typed decision model. The option probabilities are read after every loop, and every loop is trained with a proper scoring rule, so that one model serves every budget from one loop to eight in a single run. On 10,027 test decisions from 59 sources, SanSi reaches 72.0% accuracy: 13.5 points above a non-looped model of the same shape trained with the same recipe, 5.3 points above a newer non-looped model of its size, and 1.8 points below one with three times the parameters. On two depth-controlled tasks, loops extend the solvable depth beyond the depths seen in training, where the larger single-pass model fails. Used as the judge for policy optimization with reinforcement learning, without gold answers, SanSi raises the generator's F1 by 7.7 points.
Figures & tables
| Test | ||||
| Item type | Train | In-dist. | Near | Far |
| A Classification | 2,400 | 321 | – | 768 |
| B Multi-step reasoning | 5,500 | 800 | 991 | 1,281 |
| C Uncertain evidence | 2,300 | 478 | 640 | 440 |
| D Long documents | 1,000 | 357 | – | 544 |
| E Sentence pairs | 1,600 | 160 | 240 | 968 |
| Accuracy (%) read at loop | ECE at loop | |||||
| Training | 1 | 2 | 4 | 8 | 4 | 8 |
| SanSi : every loop, cross-entropy + Brier, | 58.4 0.7 | 66.9 0.6 | 71.6 0.5 | 72.0 0.7 | .085 .009 | .093 .012 |
| Loops that carry the loss | ||||||
| Loops 1, 2, 4, 8 only | 59.2 0.8 | 67.0 0.3 | 71.0 0.2 | 71.2 0.2 | .090 .013 | .098 .014 |
| Last loop only | 35.6 4.7 | 51.9 7.9 | 67.9 1.5 | 70.9 0.4 | .069 .017 | .073 .009 |
| Every loop, | 59.7 0.5 | 67.6 0.7 | 70.8 0.4 | 69.6 0.1 | .094 .005 | .106 .005 |
Appendix figures & tables44 assets
Supplementary material from the paper’s appendix.
Appendix
| Item type | Sources (selection) | Example decision | Train | Test |
|---|---|---|---|---|
| A Classification | AG News, DBpedia, Yelp ( Zhang et al., 2015 ) , IMDB ( Maas et al., 2011 ) , SST-5 ( Socher et al., 2013 ) , TREC ( Li and Roth, 2002 ) , Banking77 ( Casanueva et al., 2020 ) , TweetEval ( Barbieri et al., 2020 ) | “Meh. I was unimpressed.” Would this reviewer recommend the business? (A) no (B) yes A | 2,400 | 1,089 |
| B Multi-step reasoning | CLUTRR ( Sinha et al., 2019 ) , ProofWriter ( Tafjord et al., 2021 ) , MuSiQue ( Trivedi et al., 2022 ) , FOLIO ( Han et al., 2022 ) , BBH ( Suzgun et al., 2023 ) | “No homework is fun. Some reading is homework.” Is the statement “Some reading is fun.” true, false, or unknown? (A) true (B) false (C) unknown C | 5,500 | 3,072 |
| C Uncertain evidence | ChaosNLI ( Nie et al., 2020b ) , Sys1Cal ( Porcedda, 2026 ) , question pairs with the evidence removed (from MuSiQue) | “A hockey fight.” Hypothesis: “fighting on the ice”. (A) entailment (B) neutral (C) contradiction the annotators’ label distribution | 2,300 | 1,558 |
| D Long documents | HotpotQA ( Yang et al., 2018 ) , ContractNLI ( Koreeda and Manning, 2021 ) , QuALITY ( Pang et al., 2022 ) | [two encyclopedia paragraphs] Are Anja Salomonowitz and Rod Lurie both directors? (A) no (B) yes B | 1,000 | 901 |
| E Sentence pairs | MNLI ( Williams et al., 2018 ) , BoolQ ( Clark et al., 2019 ) , ANLI ( Nie et al., 2020a ) , WANLI ( Liu et al., 2022 ) , PAWS ( Zhang et al., 2019 ) , QNLI ( Wang et al., 2019 ) | “Revco was acquired in 1997 by CVS.” Does this sentence mean the same thing: “Revco was subsequently acquired by CVS in 1997.” (A) no (B) yes B | 1,600 | 1,368 |
| F Knowledge | ARC ( Clark et al., 2018 ) , CommonsenseQA ( Talmor et al., 2019 ) , MMLU ( Hendrycks et al., 2021 ) , SciQ ( Welbl et al., 2017 ) | Coal is formed from (A) seas that have evaporated … (D) plant remains decomposed under pressure D | – | 1,808 |
| Item type | Seen in training (training / in-distribution test items) | Near transfer | Far transfer |
|---|---|---|---|
| A Classification | 2,400 / 321 : AG News † 300/41; Amazon † 300/40; Banking77 † 300/40; DBpedia † 300/40; IMDB † 300/40; SST-5 † 300/40; TREC † 300/40; Yelp † 300/40 | – | 768 : Emotion 240; TweetEval 240; Yahoo Answers 160; Emotion ‡ 64; Offensive tweets ‡ 64 |
| B Multi-step reasoning | 5,500 / 800 : ProofWriter 1,500/240; CLUTRR, 2–4 hops 1,200/160; Kev rules † 1,110/120; MuSiQue, 2–3 hops 1,000/160; Kev policies † 690/120 | 991 : CLUTRR, 5–10 hops 480; ProofWriter, paraphrased 319; MuSiQue, 4 hops 192 | 1,281 : BBH 480; HoVer 360; FOLIO 320; Kev policies, held out ‡ 64; Kev rules, held out ‡ 57 |
| C Uncertain evidence | 2,300 / 478 : MuSiQue pairs, 2–3 hops 1,000/158; SQuAD 2.0 800/240; ChaosNLI (SNLI, NLI) 500/80 | 640 : ChaosNLI (MNLI) 400; MuSiQue pairs, 4 hops 240 | 440 : Sys1Cal 292; Kev unknowable pairs ‡ 148 |
| D Long documents | 1,000 / 357 : MuSiQue with distractors 600/237; HotpotQA 400/120 | – | 544 : ContractNLI 240; QuALITY 240; Kev buried evidence ‡ 64 |
| E Sentence pairs | 1,600 / 160 : BoolQ † 800/80; MNLI † 800/80 | 240 : MNLI, mismatched genres 240 | 968 : ANLI 240; PAWS 200; QNLI 200; WANLI 200; PAWS ‡ 64; QNLI ‡ 64 |
| F Knowledge | – | – | 1,808 : MMLU 480; ARC-Challenge 320; CommonsenseQA 240; MMLU-Pro 240; MMLU-Pro ‡ 160; ARC-Easy 120; SciQ 120; MMLU ‡ 64; SciQ ‡ 64 |
| Backbone | Trained | Training | Test | |
| Model | params | params | (GPU-min) | (GPU-s) |
| SmolLM2-1.7B | 1.71B | 72.6M | 39 † | 382 |
| Ouro-1.4B, one loop | 1.43B | 60.8M | 33 | 334 |
| SanSi , | 1.43B | 60.8M | 313 | 2,571 |
| SanSi -2.6B, | 2.67B | 121.4M | 558 § | 4,963 |
| Qwen3.5-2B | 1.88B | 67.5M | 48 ‡ | 387 |
| Difference (points) | All | In-dist. | Near | Far | JevBench |
|---|---|---|---|---|---|
| SanSi L8 SmolLM2-1.7B | [ , ] | [ , ] | [ , ] | [ , ] | [ , ] |
| SanSi L8 Ouro one loop | [ , ] | [ , ] | [ , ] | [ , ] | [ , ] |
| Ouro one loop SmolLM2-1.7B | [ , ] | [ , ] | [ , ] | [ , ] | [ , ] |
| SanSi L1 Ouro one loop | [ , ] | [ , ] | [ , ] | [ , ] | [ , ] |
| SanSi (L4) SmolLM2-1.7B | [ , ] | [ , ] | [ , ] | [ , ] | [ , ] |
| Qwen3.5-4B SanSi (L4) | [ , ] | [ , ] | [ , ] | [ , ] | [ , ] |
| Accuracy (%) | ECE | Confidence | AUROC | Evid. | Hard | Answers | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Loop | All | In-dist. | Near | Far | JevB. | In-dist. | Near | Far | right | wrong | r/w | AUROC | (%) | changed (%) |
| 1 | 58.4 | 76.3 | 61.1 | 50.8 | 64.5 | .020 | .053 | .156 | .769 | .567 | .760 | .835 | 27.1 | – |
| 0.7 | 0.9 | 1.3 | 0.5 | 0.7 | .006 | .008 | .019 | .020 | .026 | .005 | .011 | 2.1 | ||
| 2 | 66.9 | 82.8 | 64.2 | 61.9 | 67.0 | .044 | .110 | .095 | .822 | .619 | .777 | .902 | 19.1 | 29.2 |
| 0.6 | 0.4 | 1.2 | 0.8 | 1.5 | .005 | .013 | .016 | .014 | .015 | .008 | .002 | 1.4 | 0.3 | |
| 3 | 70.4 | 85.6 | 67.0 | 65.9 | 71.1 | .047 | .119 | .086 | .847 | .634 | .788 | .921 | 18.8 | 14.9 |
| Every loop trained ( SanSi ) | Last loop only | |||||
|---|---|---|---|---|---|---|
| Loops | changed | wrong right | right wrong | changed | wrong right | right wrong |
| 1 2 | 29.2 0.3 | 15.3 0.8 | 6.8 0.2 | 58.4 8.9 | 29.2 5.5 | 11.7 2.2 |
| 2 3 | 14.9 0.4 | 7.6 0.2 | 3.9 0.1 | 31.5 13.3 | 17.9 7.9 | 6.0 2.1 |
| 3 4 | 7.8 0.4 | 3.6 0.4 | 2.4 0.1 | 15.4 4.9 | 8.5 3.2 | 3.4 0.7 |
| 4 5 | 4.5 0.3 | 1.9 0.2 | 1.7 0.1 | 8.2 1.8 | 4.0 0.8 | 2.3 0.5 |
| 5 6 | 3.3 0.3 | 1.4 0.2 | 1.1 0.2 | 5.3 0.8 | 2.5 0.5 | 1.6 0.2 |
| Since loop 1 (%) | Settled | Up to loop 8 (%) | |||||
|---|---|---|---|---|---|---|---|
| Loop | changed | fixed | broken | by (%) | already final | to be fixed | to be broken |
| 1 | 0.0 0.0 | 0.0 0.0 | 0.0 0.0 | 56.7 0.2 | 64.5 0.8 | 21.0 0.3 | 7.3 0.3 |
| 2 | 29.2 0.3 | 15.3 0.8 | 6.8 0.2 | 74.2 1.2 | 78.7 0.7 | 11.1 0.7 | 5.9 0.2 |
| 3 | 33.2 0.4 | 19.0 0.8 | 6.9 0.1 | 83.6 1.1 | 85.8 0.8 | 6.4 0.7 | 4.8 0.2 |
| 4 | 34.6 0.4 | 20.5 0.5 | 7.1 0.1 | 89.1 0.9 | 90.2 0.9 | 4.0 0.5 | 3.6 0.4 |
| 5 | 35.0 0.5 | 20.8 0.3 | 7.2 0.2 | 92.5 0.8 | 93.1 0.7 | 2.7 0.4 | 2.6 0.3 |
| Settles at loop | Items (%) | Right at loop 8 (%) | Confidence at loop 8 |
|---|---|---|---|
| 1 | 56.7 0.2 | 82.2 1.1 | .894 .014 |
| 2 | 17.5 0.9 | 68.7 1.2 | .795 .020 |
| 3 | 9.4 0.2 | 60.1 0.5 | .714 .017 |
| 4 | 5.5 0.3 | 50.7 3.3 | .645 .031 |
| 5 | 3.4 0.2 | 43.6 4.3 | .581 .032 |
| 6 | 2.8 0.3 | 44.1 2.5 | .540 .032 |
| Single-pass | Items | Settles at loop | Never changes | Accuracy (%) at loop | |
|---|---|---|---|---|---|
| models right | (%) | (mean) | (%) | 1 | 8 |
| 3 | 45.5 0.1 | 1.41 0.02 | 80.8 0.6 | 86.3 0.9 | 95.0 0.2 |
| 2 | 22.9 0.4 | 2.39 0.06 | 41.4 1.5 | 52.3 1.6 | 77.4 1.5 |
| 1 | 16.0 0.3 | 3.01 0.14 | 27.5 1.9 | 28.0 0.6 | 47.1 2.0 |
| 0 | 15.6 0.3 | 2.80 0.10 | 39.0 2.1 | 13.8 0.1 | 19.8 2.3 |
| Share | Mean confidence at loop | ||||
|---|---|---|---|---|---|
| Items (single gold answer) | (%) | 1 | 2 | 4 | 8 |
| Never changed, right at loop 8 | 48.0 0.5 | .805 .019 | .883 .013 | .918 .013 | .924 .011 |
| Never changed, wrong at loop 8 | 9.5 0.7 | .669 .031 | .756 .022 | .779 .024 | .782 .022 |
| Changed at least once, right at loop 8 | 24.7 0.4 | .567 .024 | .647 .015 | .734 .018 | .763 .017 |
| Changed at least once, wrong at loop 8 | 17.7 0.3 | .530 .023 | .559 .020 | .575 .020 | .595 .021 |
| Difference (first second) | First | Second | Difference [95% interval] |
|---|---|---|---|
| ECE (answerable items) | |||
| SanSi loop 3 loop 1 | 0.082 | 0.105 | [ , ] |
| SanSi loop 8 loop 3 | 0.093 | 0.082 | [ , ] |
| SanSi loop 1 Ouro one loop | 0.105 | 0.137 | [ , ] |
| SanSi loop 8 Ouro one loop | 0.093 | 0.137 | [ , ] |
| SanSi loop 8 SmolLM2-1.7B | 0.093 | 0.069 | [ , ] |
| Loop | Accuracy (%) | Mean confidence | Confidence accuracy | ECE |
|---|---|---|---|---|
| 1 | 57.9 0.7 | .683 .022 | .105 .020 | |
| 2 | 66.3 0.6 | .750 .015 | .088 .009 | |
| 3 | 70.0 0.6 | .780 .016 | .082 .010 | |
| 4 | 71.2 0.6 | .793 .016 | .085 .009 | |
| 5 | 71.4 0.7 | .798 .016 | .087 .008 | |
| 6 | 71.7 0.7 | .802 .016 | .089 .010 |
| Item | Answer | |
| Liar chains (options: yes, no) | ||
| 1 | Lee is honest. Wes is honest. According to Flo, Lee tells the truth. According to Abe, Wes lies. Does Abe tell the truth? | no |
| 2 | Uma always tells the truth. Seth always tells the truth. Fred says that Seth lies. Kurt says Fred is a liar. Bert says that Uma lies. Ben says Bert is a liar. Is Ben telling the truth? | yes |
| 4 | Omar is honest. Ege always lies. According to Bert, Omar tells the truth. Yves says that Eli lies. Ben says that Bert lies. Liv says that Nate lies. Eli says that Liv tells the truth. Jon says Vera is honest. Vera says that Ben lies. According to Nate, Ege tells the truth. Is Yves telling the truth? | no |
| 8 | Omar is a liar. Seth is a liar. Vera says Iris is a liar. Bea says Flo is a liar. According to Eli, Wade tells the truth. According to Yul, Ida tells the truth. Nia says that Nate lies. According to Bert, Omar lies. Ida says that Nia lies. Iris says Bert is honest. According to Wade, Meg lies. Nate says Vera is honest. According to Meg, Bo tells the truth. Ben says that Seth tells the truth. According to Kim, Bea tells the truth. Flo says Ben is a liar. According to Ola, Yul lies. Bo says that Kim lies. Is Eli telling the truth? | no |
| Object swaps (options: the five people) | ||
| Holds | Accuracy (%) | Difference to Qwen3.5-4B (points) | |||
|---|---|---|---|---|---|
| Model | depth | ||||
| Liar chains (chance 50%) | |||||
| SanSi , loop 1 | 3 | 69.0 | 49.9 | [ , ] | [ , ] |
| SanSi , loop 2 | 6 | 88.0 | 51.6 | [ , ] | [ , ] |
| SanSi , loop 4 | 11 | 96.6 | 70.9 | [ , ] | [ , ] |
| SanSi , loop 8 | 11 | 97.0 | 72.6 | [ , ] | [ , ] |
| Depths seen in training | Unseen depths | |||||||||||||||
| 1 | 2 | 3 | 4 | 5 | 6 | 7 | 8 | 9 | 10 | 11 | 12 | 13 | 14 | 15 | 16 | |
| Liar chains (chance 50%) | ||||||||||||||||
| SanSi , loop 1 | 100.0 | 98.6 | 79.2 | 60.0 | 59.4 | 52.8 | 50.6 | 51.4 | 48.6 | 50.6 | 51.7 | 50.6 | 50.8 | 48.3 | 49.2 | 49.2 |
| 0.0 | 1.3 | 9.8 | 11.0 | 3.4 | 8.0 | 7.1 | 1.7 | 3.8 | 1.7 | 4.4 | 1.3 | 2.5 | 1.4 | 2.2 | 1.4 | |
| SanSi , loop 2 | 100.0 | 100.0 | 99.7 | 96.7 | 91.4 | 80.8 | 69.7 | 65.6 | 55.8 | 51.7 | 51.1 | 52.5 | 52.5 | 50.6 | 51.1 | 47.5 |
| 0.0 | 0.0 | 0.5 | 2.2 | 2.7 | 6.0 | 7.7 | 5.1 | 3.6 | 3.6 | 4.3 | 0.8 | 3.8 | 5.7 | 1.9 | 3.0 | |
| 1 | 2 | 3 | 4 | 5 | 6 | 7 | 8 | 9 | 10 | 11 | 12 | 13 | 14 | 15 | 16 | |
| Liar chains | ||||||||||||||||
| Settles at loop | 1.0 | 1.0 | 1.2 | 1.5 | 1.6 | 1.9 | 2.1 | 2.3 | 2.7 | 2.9 | 3.1 | 3.4 | 3.9 | 4.0 | 4.2 | 4.2 |
| Changed, loop 1 to 8 (%) | 0 | 1 | 21 | 40 | 41 | 47 | 50 | 52 | 54 | 52 | 52 | 54 | 55 | 56 | 53 | 52 |
| Object swaps | ||||||||||||||||
| Settles at loop | 1.0 | 1.1 | 1.4 | 1.9 | 2.0 | 2.3 | 2.7 | 3.0 | 3.0 | 3.3 | 3.5 | 3.6 | 4.0 | 4.1 | 4.0 | 4.5 |
| Changed, loop 1 to 8 (%) | 1 | 3 | 32 | 53 | 51 | 65 | 59 | 70 | 64 | 68 | 66 | 66 | 68 | 64 | 64 | 64 |
| 1 | 2 | 3 | 4 | 5 | 6 | 7 | 8 | 9 | 10 | 11 | 12 | 13 | 14 | 15 | 16 | |
| SmolLM2-1.7B, recipe of Section 3 (2,000 steps) | ||||||||||||||||
| Seed 0 | 50.0 | 50.0 | 50.0 | 50.0 | 50.0 | 50.0 | 50.0 | 50.0 | 50.0 | 50.0 | 50.0 | 50.0 | 50.0 | 50.0 | 50.0 | 50.0 |
| Seed 1 | 50.0 | 50.0 | 50.0 | 50.0 | 50.0 | 50.0 | 50.0 | 50.0 | 50.0 | 50.0 | 50.0 | 50.0 | 50.0 | 50.0 | 50.0 | 50.0 |
| SmolLM2-1.7B, half the learning rate and 4,000 steps | ||||||||||||||||
| Seed 0 | 50.0 | 49.2 | 54.2 | 52.5 | 48.3 | 49.2 | 57.5 | 53.3 | 49.2 | 52.5 | 46.7 | 55.0 | 55.8 | 50.0 | 46.7 | 45.0 |
| Seed 1 | 51.7 | 52.5 | 57.5 | 54.2 | 52.5 | 50.8 | 59.2 | 50.0 | 46.7 | 46.7 | 54.2 | 54.2 | 49.2 | 50.0 | 46.7 | 52.5 |
| 1 | 2 | 3 | 4 | 5 | 6 | 7 | 8 | 9 | 10 | 11 | 12 | 13 | 14 | 15 | 16 | |
| Liar chains (chance 50%) | ||||||||||||||||
| Accuracy (%) | 100.0 | 86.4 | 62.5 | 56.9 | 58.6 | 44.7 | 48.6 | 47.8 | 44.7 | 48.3 | 51.1 | 53.3 | 47.5 | 43.9 | 42.5 | 51.1 |
| Mean confidence | .99 | .82 | .69 | .66 | .63 | .63 | .61 | .62 | .62 | .60 | .61 | .61 | .60 | .59 | .60 | .60 |
| Object swaps (chance 20%) | ||||||||||||||||
| Accuracy (%) | 77.5 | 42.5 | 23.6 | 20.6 | 24.2 | 25.0 | 19.4 | 21.4 | 18.6 | 20.3 | 26.9 | 24.2 | 21.4 | 31.1 | 23.3 | 27.8 |
| Mean confidence | .92 | .69 | .53 | .47 | .42 | .39 | .37 | .36 | .32 | .33 | .31 | .31 | .31 | .29 | .29 | .29 |
| Accuracy (%) at loop | Accuracy at loop 8 | ECE | Hard | Evid. | AUROC | |||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Training | 1 | 2 | 4 | 8 | In | Near | Far | all | near | ans. (%) | AUROC | r/w |
| Loss (cross-entropy + Brier) | 58.4 | 66.9 | 71.6 | 72.0 | 86.6 | 67.7 | 68.0 | .093 | .149 | 17.5 | .935 | .795 |
| 0.7 | 0.6 | 0.5 | 0.7 | 0.6 | 2.7 | 0.1 | .012 | .031 | 1.3 | .004 | .005 | |
| Reinforcement learning | 58.1 | 66.5 | 71.3 | 71.7 | 86.1 | 66.8 | 67.9 | .076 | .129 | 21.6 | .919 | .793 |
| 0.2 | 0.4 | 0.4 | 0.1 | 0.1 | 2.7 | 0.7 | .015 | .047 | 4.5 | .008 | .002 | |
| Difference (first second) | First | Second | Difference [95% interval] |
|---|---|---|---|
| Accuracy (%) | |||
| SanSi : loop 8 loop 4 | 72.0 | 71.6 | [ , ] |
| SanSi : loop 12 loop 8 | 71.2 | 72.0 | [ , ] |
| SanSi : loop 16 loop 8 | 70.1 | 72.0 | [ , ] |
| Four-loop model: loop 8 loop 4 | 69.6 | 70.8 | [ , ] |
| Four-loop model SanSi , both at loop 4 | 70.8 | 71.6 | [ , ] |
| SanSi (8 trained loops) | Trained with 4 loops | |||
|---|---|---|---|---|
| Loop | Acc. (%) | ECE | Acc. (%) | ECE |
| 1 | 58.4 0.7 | .105 .020 | 59.7 0.5 | .123 .010 |
| 2 | 66.9 0.6 | .088 .009 | 67.6 0.7 | .096 .005 |
| 3 | 70.4 0.6 | .082 .010 | 70.4 0.5 | .093 .007 |
| 4 | 71.6 0.5 | .085 .009 | 70.8 0.4 | .094 .005 |
| 5 | 71.9 0.7 | .087 .008 | 70.9 0.4 † | .103 .009 |
| Accuracy (%) at loop | |||
|---|---|---|---|
| Model | 1 | 4 | 8 |
| SmolLM2-1.7B, not fine-tuned | 38.8 | – | – |
| + loop as in Ouro | 38.8 | 32.0 | 31.5 |
| SmolLM2-1.7B, fine-tuned | 58.4 0.7 | – | – |
| + loop as in Ouro | 35.4 4.2 | 32.5 0.6 | 33.5 0.4 |
| + loop through a linear map | 58.1 1.1 | 58.3 1.0 | 58.3 1.0 |
| Accuracy (%) | ECE | Confidence | AUROC | Evid. | Hard | Answers | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Loop | All | In-dist. | Near | Far | JevB. | In-dist. | Near | Far | right | wrong | r/w | AUROC | (%) | changed (%) |
| 1 | 62.4 | 78.8 | 61.9 | 56.6 | 60.6 | .032 | .081 | .118 | .789 | .596 | .756 | .839 | 31.5 | – |
| 0.4 | 0.5 | 1.7 | 0.8 | 0.7 | .009 | .022 | .014 | .011 | .014 | .004 | .006 | 3.5 | ||
| 2 | 71.8 | 85.2 | 70.0 | 67.6 | 70.7 | .049 | .103 | .078 | .852 | .650 | .780 | .914 | 22.6 | 27.5 |
| 0.7 | 0.2 | 3.0 | 0.2 | 1.7 | .002 | .023 | .010 | .008 | .014 | .003 | .007 | 3.0 | 0.9 | |
| 3 | 74.7 | 87.8 | 72.5 | 70.6 | 74.6 | .049 | .117 | .073 | .873 | .660 | .795 | .935 | 18.1 | 12.2 |
| Acc. | ECE | |||
| (%) | all | near | far | |
| All test items; no calibration data | ||||
| Loop 8 | 72.0 0.7 | .093 .012 | .149 .031 | .092 .019 |
| Mean of loops 1–8 | 71.8 0.7 | .044 .009 | .084 .028 | .045 .010 |
| Temperature fitted on in-distribution items | ||||
| Loop 8 | 70.2 0.7 | .098 .013 | .151 .032 | .092 .019 |
| Reward | F1 after training | Change in F1 [95% interval] | Reward AUROC, first 20 steps |
|---|---|---|---|
| Not trained | 39.5 | – | – |
| SanSi read at loop 1 | 29.1 2.4 | [ , ] | .779 .050 |
| SanSi read at loop 2 | 39.5 1.8 | [ , ] | .857 .030 |
| SanSi read at loop 4 | 45.8 4.2 | [ , ] | .914 .032 |
| SanSi read at loop 8 | 47.3 1.1 | [ , ] | .905 .025 |
| Loop 8 loop 4 | – | [ , ] | – |
| All questions | F1 by question type | Answer | ||||||
|---|---|---|---|---|---|---|---|---|
| Reward | F1 | EM | Contains | Comparison | Bridge comparison | Compositional | Inference | words |
| Not trained | 39.5 | 32.5 | 34.5 | 50.6 | 48.2 | 30.7 | 28.6 | 2.6 |
| SanSi read at loop 1 | 29.1 2.4 | 9.6 3.0 | 38.5 1.0 | 34.0 3.2 | 32.7 4.9 | 28.6 2.8 | 21.0 2.7 | 5.5 0.5 |
| SanSi read at loop 2 | 39.5 1.8 | 19.6 5.4 | 45.0 1.9 | 38.9 2.9 | 42.4 7.5 | 36.3 1.5 | 40.5 3.4 | 4.3 0.5 |
| SanSi read at loop 4 | 45.8 4.2 | 29.4 10.7 | 49.7 0.9 | 49.2 8.2 | 48.5 6.2 | 37.3 1.9 | 48.4 1.1 | 3.9 0.8 |
| SanSi read at loop 8 | 47.3 1.1 | 32.2 3.8 | 49.7 2.0 | 49.4 5.1 | 51.6 0.7 | 38.3 1.5 | 49.7 0.9 | 3.6 0.3 |
| F1 | EM | Contains | Answer words | |||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Reward, seed | 0 | 1 | 2 | 0 | 1 | 2 | 0 | 1 | 2 | 0 | 1 | 2 |
| SanSi read at loop 1 | 27.4 | 27.9 | 31.9 | 8.2 | 7.5 | 13.0 | 37.5 | 38.4 | 39.5 | 5.9 | 5.6 | 4.9 |
| SanSi read at loop 2 | 40.2 | 37.4 | 40.9 | 24.0 | 13.6 | 21.3 | 43.0 | 45.4 | 46.7 | 3.9 | 4.8 | 4.3 |
| SanSi read at loop 4 | 48.3 | 48.2 | 41.0 | 36.3 | 34.8 | 17.1 | 48.9 | 49.5 | 50.7 | 3.4 | 3.5 | 4.7 |
| SanSi read at loop 8 | 48.5 | 46.5 | 46.8 | 36.5 | 30.7 | 29.3 | 48.7 | 48.3 | 52.0 | 3.4 | 3.5 | 4.0 |
| Share of the items (%) | |
|---|---|
| SanSi wrong at loop 8 | 27.2 |
| Qwen3.5-4B wrong | 25.7 |
| SmolLM2-1.7B wrong | 41.1 |
| SanSi and Qwen3.5-4B both wrong | 18.5 |
| Only SanSi wrong | 8.7 |
| Only Qwen3.5-4B wrong | 7.2 |
| SanSi | ||||||
|---|---|---|---|---|---|---|
| Item (source) | Gold | loop 1 | loop 8 | Qwen3.5-4B | SmolLM2 | |
| Fixed by the loops | “No road is dustless. Some streets are roads.” Is “Some streets are dustless.” true, false, or unknown? (FOLIO) | unknown | true (.83) | unknown (.97) | unknown (.89) | true (.57) |
| Fixed; larger model wrong | If 30,000 is divided by 10 and then divided by 10 again, what will be the resulting number? 3 / 30 / 300 / 3,000 (MMLU) | 300 | 3 (.60) | 300 (.91) | 3,000 (.87) | 300 (.38) |
| Missing knowledge | When cold temperatures are produced in a chemical reaction, the reaction is known as … (ARC) | endothermic | exothermic (.96) | exothermic (.99) | endothermic (.97) | exothermic (.48) |
| Broken by the loops | “By 9000 BP, Europe was fully forested.” Does the sentence contain the answer to “When was Europe fully forested and recovered from the last Ice Age?” (QNLI) | yes | yes (.98) | no (.89) | no (.77) | yes (.96) |
| Shared error | “He can’t be here.” Hypothesis: “He is here.” (WANLI) | neutral | contradiction (.99) | contradiction (.96) | contradiction (.99) | contradiction (.71) |