When the Right Answer Is Missing: An Arithmetic-Dependent Rejection Bottleneck in Jev
Organizations: University of Southern California · University of Florida · Emory University
Abstract
Typed decision models such as Jev offer an efficient alternative to generative LLMs in decision-making workflows by selecting directly from predefined options. When candidate sets contain no valid answer, TypeSafe recommends including an "other" or "none-of-the-above" option to enable rejection. In this report, however, we identify an arithmetic-dependent rejection bottleneck: Jev reliably selects correct numerical answers when available but frequently accepts incorrect alternatives when they are absent despite an explicit rejection option. On paired arithmetic problems, answer-present accuracy reaches 99%, while correct rejection falls to 7%. Moreover, this gap persists across numerical magnitudes, operation depths, contextual formulations, and rejection labels, and extends to scenarios such as time calculation and capacity rounding. Yet native Boolean verification achieves 99% exact-match accuracy on the same answer-absent arithmetic cases, showing that categorical rejection can fail even when the model successfully verifies candidate correctness. Finally, we show that a simple decision threshold selected on separate development problems raises arithmetic rejection accuracy from 7% to 79% while retaining 97% answer-present accuracy, substantially mitigating the failure without retraining or additional inference.
Figures & tables
| Two-operation pattern | Present | Absent |
|---|---|---|
| 9/10 | 10/10 | |
| 9/10 | 3/10 | |
| 8/10 | 1/10 | |
| 10/10 | 1/10 | |
| 8/10 | 0/10 | |
| 10/10 | 8/10 |
| Task | Interface | Present | Absent | |
|---|---|---|---|---|
| Arithmetic | Menu | .03 | 99 97 | 7 79 |
| T/F | .74 | 61 86 | 27 90 | |
| Boolean | .50 | 97 97 | 99 99 | |
| Purchase | Menu | .03 | 94 86 | 24 79 |
| T/F | .80 | 45 73 | 17 82 | |
| Boolean | .90 | 73 87 | 76 97 |