System Switch: When Should a Fast Decision Model Stop and Think?
Organizations: Independent Researcher, Recco (Genoa), Italy
Abstract
Dual-process agents pair a fast policy with a slow deliberative model. In real-time settings the slow model usually runs continuously; in turn-based agents and robot planners it is invoked on events such as uncertainty or a detected failure. We study a fast learned actor that takes every decision and hands control to a reasoning vision-language model only when a gate opens, while the game keeps running. We use closed-loop Doom and the new open "System One" typed-decision models, served through a common llama.cpp interface. On 900 held-out questions, (i) zero-shot decision models from 0.15B to 9B parameters choose to collect items 1.6-1.8 times more often than chance among their errors, in any option order, although the order changes some models' accuracy; (ii) accuracy, calibration and sensitivity (how well confidence separates right from wrong answers) are distinct: models of similar accuracy differ widely in AUROC, and the confidence of the most sensitive one tracks which kinds of situation it fails, not which answers are wrong; (iii) offline, deferring the least confident 30% of decisions to a reasoning model gains over random deferral in proportion to the actor's AUROC (rank correlation 0.87); with actor and rate chosen on held-out games the gain is +0.13 [0.08, 0.18] with doomLaya's option order and +0.08 [0.02, 0.14] with shuffled options, and reasoning carries about half of it; (iv) in closed loop (33 games, three seeds) no variant reaches the exit. Committing to plans, the reasoner's or a fixed explore rule's, opens more doors and makes an actor that stands still play; with the rule the agent dies more often. Told that some doors need keys, the reasoner takes ordinary doors for locked ones, which the state cannot tell apart; without that knowledge it goes back to collecting. We release code, prompts, data and logs.
Figures & tables
| command | |||||
| model | params | fixed order | shuffled | weapon | collect ( chance) |
| Kev | 9B | 0.60 [0.49, 0.69] | 0.50 | 1.00 | 92% (1.7) |
| Strands Decider | 2B | 0.59 [0.49, 0.68] | 0.38 | 1.00 | 97% (1.8) |
| Laya, typed-decisions | 0.4B | 0.59 [0.48, 0.68] | 0.56 | 0.42 | 91% (1.6) |
| Laya, base | 0.4B | 0.56 [0.45, 0.66] | 0.49 | 0.38 | 83% (1.7) |
| Clef-Flash | 9B | 0.54 [0.43, 0.63] | 0.54 | 0.60 | 92% (1.7) |
| AUROC, fixed order | AUROC | ||||
| actor | accuracy | ECE | overall | within category | shuffled |
| Laya, base | 0.56 | 0.16 | 0.89 [0.85, 0.92] | 0.52 [0.40, 0.63] | 0.72 |
| Kev 0.8B | 0.29 | 0.14 | 0.83 [0.76, 0.88] | 0.75 [0.63, 0.85] | 0.77 |
| Laya, typed-decisions | 0.59 | 0.29 | 0.77 [0.70, 0.83] | 0.65 [0.42, 0.81] | 0.69 |
| Kev 4B | 0.40 | 0.10 | 0.70 [0.58, 0.80] | 0.80 [0.68, 0.89] | 0.66 |
| Clef-Flash 9B | 0.54 | 0.10 | 0.68 [0.59, 0.76] | 0.68 [0.58, 0.78] | 0.68 |
| deferred to the thinker | confidence random | ||||
| actor (AUROC) | alone | 10% | 30% | 50% | at 30% |
| thinker: Qwen3.6 with reasoning (alone: 0.69 [0.60, 0.76]) | |||||
| Laya, base (0.89) | 0.56 | 0.61 | 0.75 | 0.76 | +0.15 [0.11, 0.19] |
| Laya-v3, trained (0.87) | 0.59 | 0.62 | 0.72 | 0.82 | +0.11 [0.06, 0.18] |
| Laya, typed-decisions (0.77) | 0.59 | 0.60 | 0.67 | 0.72 | +0.05 [0.02, 0.08] |
| Kev 9B (0.61) | 0.60 | 0.60 | 0.62 | 0.63 | 0.01 [ 0.04, 0.01] |
| variant | kills | deaths | cells | doors | closest | thoughts | guided | thinker’s top choice |
| actor: Laya, typed-decisions | ||||||||
| off | 11.0 | 0.7 | 77 | 5.3 | 12.6 m | – | – | – |
| bare thinker | 9.3 | 0.7 | 61 | 5.0 | 12.8 m | 14.3 | 7% | collect (49%) |
| commitment, gate: unsure | 12.0 | 2.0 | 66 | 8.0 | 12.9 m | 7.0 | 28% | explore (52%) |
| commitment, gate: progress | 14.3 | 1.7 | 83 | 8.7 | 12.6 m | 4.0 | 18% | explore (67%) |
| commitment, gate: all | 14.3 | 1.3 | 83 | 8.0 | 9.5 m | 7.0 | 26% | explore (57%) |