This paper asks whether a bounded neural architecture can exhibit a meaningful division of labor between intuition and deliberation on a classic 64-item syllogistic reasoning benchmark. More broadly, the benchmark is relevant to ongoing debates about world models and multi-stage reasoning in AI. It provides a controlled setting for testing whether a learned system can develop structured internal computation rather than only one-shot associative prediction. Experiment 1 evaluates a direct neural baseline for predicting full 9-way human response distributions under 5-fold cross-validation. Experiment 2 introduces a bounded dual-path architecture with separate intuition and deliberation pathways, motivated by computational mental-model theory (Khemlani & Johnson-Laird, 2022). Under cross-validation, bounded intuition reaches an aggregate correlation of r = 0.7272, whereas bounded deliberation reaches r = 0.8152, and the deliberation advantage is significant across folds (p = 0.0101). The largest held-out gains occur for NVC, Eca, and Oca, suggesting improved handling of rejection responses and c-a conclusions. A canonical 80:20 interpretability run and a five-seed stability sweep further indicate that the deliberation pathway develops sparse, differentiated internal structure, including an Oac-leaning state, a dominant workhorse state, and several weakly used or unused states whose exact indices vary across runs. These findings are consistent with reasoning-like internal organization under bounded conditions, while stopping short of any claim that the model reproduces full sequential processes of model construction, counterexample search, and conclusion revision.
Figures & tables
Label
Reading
Aac
All a are c
Eac
No a are c
Iac
Some a are c
Oac
Some a are not c
Aca
All c are a
Eca
No c are a
Table 1: Response categories used in the human target distribution.
Model
Role in paper
Input
Hidden / state size
Output
Parameters
Direct MLP
Experiment 1 baseline
29
64
9-way softmax distribution
2,505
Intuition pathway
Experiment 2 first-pass system
29
4
9-way softmax distribution
165
Deliberation pathway
Experiment 2 second-stage system
29
24
9-way softmax distribution
470
Experiment 2 total
Combined bounded model
29
4 + 24
2 x 9-way softmax distributions
635
Table 2: Summary of the model variants used in the paper.
Figure 1: Architecture of the bounded intuition-and-deliberation model used in Experiment 2. The intuition pathway and the deliberation pathway receive the same structured syllogism input through separate encoders, and the deliberation pathway computes through five candidate deliberative states combined by a learned gate.
Model
Aggregate correlation
95% bootstrap CI
Aggregate RMSE
Aggregate MAE
Mean fold correlation
Fold SD
Experiment 1 direct MLP
0.7105
[0.6309, 0.7825]
0.1301
0.0675
0.7140
0.0980
Experiment 2 intuition
0.7272
[0.6649, 0.7868]
0.1156
0.0674
0.7212
0.0838
Experiment 2 deliberation
0.8152
[0.7608, 0.8634]
0.1029
0.0536
0.8142
0.0710
Table 3: Main 5-fold cross-validation results. The deliberation pathway shows the strongest overall held-out performance.
Response type
Direct MLP
Intuition
Deliberation
Deliberation - Intuition
Aac
0.9718
0.7353
0.8450
+0.1097
Eac
0.7816
0.8259
0.8156
-0.0103
Iac
0.7741
0.6894
0.8441
+0.1547
Oac
0.8033
0.6336
0.7050
+0.0714
Aca
-0.0170
0.1682
0.2169
+0.0488
Eca
0.4050
0.5203
0.7790
+0.2587
Table 4: Aggregate per-response-type correlations under 5-fold cross-validation.
Figure 2: Representative held-out performance for the direct baseline in Experiment 1. The three panels show the human response matrix, the model prediction matrix, and the signed difference matrix on the test split.
Comparison
t
p
Interpretation
Deliberation vs intuition
4.5914
0.0101
Significant deliberation advantage
Deliberation vs direct MLP
1.5678
0.1920
Numerical advantage, not significant
Intuition vs direct MLP
0.1047
0.9217
No reliable difference
Table 5: Paired 5-fold statistical comparisons using fold-wise correlations.
Fold
Experiment 1 direct MLP
Experiment 2 intuition
Experiment 2 deliberation
1
0.8523
0.6599
0.7313
2
0.6223
0.6271
0.7478
3
0.5859
0.8284
0.8627
4
0.7350
0.7391
0.8256
5
0.7745
0.7517
0.9036
Table 6: Fold-wise correlations used in the paired statistical comparisons.
Figure 3: Intuition-versus-deliberation performance on the canonical single split used for interpretability. Deliberation improves both correlation and error relative to intuition.
Figure 4: Held-out response-distribution predictions from the bounded intuition pathway in Experiment 2.
Figure 5: Held-out response-distribution predictions from the deliberation pathway in Experiment 2. Relative to intuition, deliberation more closely tracks the human response matrix on the same held-out items.
Seed
Test intuition r
Test deliberation r
Deliberation gain
Dominant test state
Dead states
Oac-leaning state
Strongest ablation state
1
0.7051
0.8739
+0.1688
state 2
state 3, state 5
state 3
state 2
2
0.7291
0.8072
+0.0781
state 3
state 5
state 3
state 1
3
0.7606
0.8141
+0.0535
state 1
state 2, state 4, state 5
state 5
state 5
4
0.6983
0.8061
+0.1077
state 1
state 2, state 3
state 1
state 1
5
0.8261
0.8372
+0.0111
state 4
state 2, state 3, state 5
state 4
state 4
Table 7: Summary of the five-seed stability sweep over repeated 80:20 runs.
State
Train gate winners
Test gate winners
State 1
13
1
State 2
6
1
State 3
0
0
State 4
22
11
State 5
10
0
Table 8: Gate-winner counts for the five deliberative states in the canonical 80:20 interpretability split.
State
Observed role
Main evidence
State 1
Oac -leaning specialist for O / E -minor items
Average state distribution dominated by Oac ; removal harms Oac behavior
State 2
Universal-conclusion contributor for AA and AE valid syllogisms
Wins on AA1 , AA4 , AE1 , AE3 , and AE4 ; removal causes moderate held-out loss
State 3
Effectively unused
Never wins on train or test; ablation has almost no effect
State 4
Dominant general-purpose state for many I / O -premise and NVC -heavy items
Most frequent gate winner on train and test; strongest association with Iac , Ica , Oca , and NVC ; largest ablation effect
State 5
E -premise / EE - EA contributor
Wins on EA* , EI* , and all four EE items in training; ablation still causes large performance loss
Table 9: Summary of state specialization in the canonical interpretability split.
Ablation
Test correlation
Correlation drop
Test RMSE
RMSE increase
none
0.8475
0.0000
0.0946
0.0000
remove state 1
0.7331
0.1143
0.1310
0.0364
remove state 2
0.7384
0.1091
0.1292
0.0345
remove state 3
0.8474
0.0000
0.0946
0.0000
remove state 4
0.2642
0.5833
0.2700
0.1754
remove state 5
0.5519
0.2956
0.1796
0.0850
Table 10: Inference-time ablation results on the held-out test partition of the canonical interpretability split.