Understanding as No-Arbitrage: Bounded Dutch Books as a Definition and Training Objective for Language Models
Abstract
Does a language model merely predict tokens, or does it understand what it says? We make this question measurable by defining "understanding" through the lens of no-arbitrage. A model understands a vocabulary to a certain degree if a computationally bounded trader cannot extract guaranteed profit by betting against the model's probabilities on logically related claims (a "Dutch book"). We establish three theoretical results: first, because full logical coherence is computationally intractable, understanding is inherently graded, not absolute. Second, we prove that the exact optimum of standard next-token prediction is inherently incoherent across different question formats; the flaw lies in the training objective, not the architecture. Third, we show that uncertainty accumulates predictably along reasoning chains, making unjustified overconfidence an arbitrage opportunity in itself. To address this, we introduce Arbitr, a training framework where an adversarial trader penalizes the model for logical inconsistencies, paired with a calibration anchor to prevent uninformative collapse. Across five pre-registered experiments on Qwen2.5 and Phi-3.5 models, we demonstrate that standard models are highly exploitable across different phrasings. Arbitr reduces this exploitability by orders of magnitude without sacrificing task accuracy, and the effect successfully transfers to unseen logical patterns and new model families. Crucially, we uncover a scaling illusion: at 7B parameters, near-zero measured incoherence often coincides with extreme, unjustified confidence. We conclude that while Arbitr enforces rigorous logical consistency, coherence is a necessary condition for knowledge, but not a sufficient one
Figures & tables
| Qwen2.5-1.5B | Qwen2.5-3B | Qwen2.5-7B | SmolLM3-3B | |
|---|---|---|---|---|
| Primary median (pre-reg.) | 0.024 | 0.000 | 0.0001 | 0.019 |
| Negation channel, F1 / F2 (median) | 0.472 / 0.165 | 0.177 / 0.001 | 0.038 / 0.105 | 0.052 / 0.117 |
| Conditionals (MP, exploratory) | 0.345 | 0.500 | 0.495 | 0.439 |
| Cross-format book F1 F2: mean profit | 0.151 | 0.021 | 0.064 | 0.070 |
| share of propositions with profit | 25.0% | 4.0% | 13.3% | 7.5% |
| median by relation | capital-of | higher-than | longer-than | older-than | ball-in-box |
|---|---|---|---|---|---|
| Qwen2.5-1.5B | 0.027 | 0.249 | 0.081 | 0.075 | 0.091 |
| Qwen2.5-3B | 0.000 | 0.000 | 0.000 | 0.000 | 0.001 |
| Qwen2.5-7B | 0.000 | 0.428 | 0.482 | 0.000 | 0.007 |
| SmolLM3-3B | 0.002 | 0.115 | 0.038 | 0.062 | 0.045 |
| base | A (task only) | B (task arbitrage) | |
|---|---|---|---|
| Negation Inc, median (held-out) | 0.237 | ||
| Cross-format LOP profit, mean | 0.099 | ||
| Transitivity Inc, mean (untrained, level 3) | 0.004 | ||
| Conjunction Inc, mean (untrained, level 3) | 0.070 | ||
| Held-out fact accuracy | 0.553 | ||
| Transitive-chain QA accuracy | 0.720 |
| base | A (control, 10 seeds) | B 0.3 (main, 10 seeds) | B 1.0 (dose, 4 seeds) | |
|---|---|---|---|---|
| Negation Inc, median | 0.237 | |||
| Cross-format LOP profit, mean | 0.097 | |||
| Transitivity Inc, mean (untrained) | 0.0036 | |||
| Held-out fact accuracy | 0.553 | |||
| Transitive-chain QA accuracy | 0.730 |