Typed decision models such as Jev offer an efficient alternative to generative LLMs in decision-making workflows by selecting directly from predefined options. When candidate sets contain no valid answer, TypeSafe recommends including an "other" or "none-of-the-above" option to enable rejection. In this report, however, we identify an arithmetic-dependent rejection bottleneck: Jev reliably selects correct numerical answers when available but frequently accepts incorrect alternatives when they are absent despite an explicit rejection option. On paired arithmetic problems, answer-present accuracy reaches 99%, while correct rejection falls to 7%. Moreover, this gap persists across numerical magnitudes, operation depths, contextual formulations, and rejection labels, and extends to scenarios such as time calculation and capacity rounding. Yet native Boolean verification achieves 99% exact-match accuracy on the same answer-absent arithmetic cases, showing that categorical rejection can fail even when the model successfully verifies candidate correctness. Finally, we show that a simple decision threshold selected on separate development problems raises arithmetic rejection accuracy from 7% to 79% while retaining 97% answer-present accuracy, substantially mitigating the failure without retraining or additional inference.
Figures & tables
Figure 1: Arithmetic-dependent rejection and threshold calibration. (a) Menu Choice achieves 99% answer-present accuracy but only 7% correct rejection even on simple three-operand arithmetic ( a+b−c ), while Boolean verification reaches 99% answer-absent exact match. (b) Rejection collapses after just two operations. (c) None scores retain a useful rejection signal. (d) A development-selected threshold calibration raises held-out rejection to 79% while retaining 97% answer-present accuracy, without additional inference.
Figure 2: Three formulations of the same decision. Menu Choice selects a candidate or None . Our T/F Choice ablation verifies each candidate through the categorical API; native Boolean changes only the question type. Candidate questions share one request. The oracle answer is shown for illustration and is not supplied in the arithmetic input.
Figure 3: Strong selection conceals task-dependent rejection failures. (a) Accuracy on 100 paired problems per task; P/A denotes answer present/absent. T/F and Boolean require all three judgments correct. Time and counting pool 20 screening and 80 fresh cases. (b) Present-minus-absent gaps with 95% paired-bootstrap intervals. Menu rejection deteriorates most on arithmetic and purchase totals; lookup eliminates the gap.
Figure 4: Rejection fails with simple arithmetic and depends on candidate distance. (a) Fifty paired cases per operation depth; depth zero supplies the result. (b) One hundred fresh two-operation cases per magnitude band. (c) Fifty paired inventory cases: larger distractor errors improve categorical rejection but not Boolean verification. Depth and distance vary within cases; magnitude uses independent cases.
Figure 5: A capable generative control. Qwen3.5-9B (w/ reasoning) on the same cases per family and coverage. P/A denotes present/absent. Both computational menus are solved perfectly. Token-limit failures account for 66 of 67 errors across menu and joint-verification responses, distinguishing output completion from valid incorrect decisions.
Two-operation pattern
Present
Absent
a×b×c
9/10
10/10
a×b+c
9/10
3/10
a×b−c
8/10
1/10
a+b×c
10/10
1/10
a−b×c
8/10
0/10
a÷b÷c
10/10
8/10
Table 1: Rejection varies with operator composition. Multiplication/division chains achieve higher rejection accuracy than expressions combining these operations with addition or subtraction.
Figure 6: Rejection gaps extend to other arithmetic-dependent tasks. Menu Choice results across 15 pilot families, with 20 paired cases each. Time calculation, capacity rounding, and calendar arithmetic show selection–rejection gaps, while several symbolic tasks succeed under both conditions. Time and counting show their initial pilot results; the dotted divider separates the five additional numerical tasks.
Figure 7: Score separation motivates threshold calibration. Empirical CDFs of Menu Choice rejection probabilities (left) and candidate truth scores for T/F Choice and Boolean (center/right). Answer-absent menus tend to receive higher rejection scores, while correct candidates generally receive higher truth scores than incorrect candidates. This separation motivates adjusting decision thresholds to reduce false acceptance. Gray dotted lines mark the default verification threshold; blue dashed lines mark development-selected thresholds.
Figure 8: Threshold adjustment recovers rejection on separate cases. (a) Arrows connect original menu decisions (open circles) to calibrated decisions (filled circles). Upward movement improves rejection; leftward movement loses selection accuracy. (b) Paired 95% bootstrap intervals quantify both effects. Arithmetic/purchase use 100 test cases; time/counting use 80. Thresholds are chosen only on the corresponding 20-case development sets.
Task
Interface
τ
Present
Absent
Arithmetic
Menu
.03
99 → 97
7 → 79
T/F
.74
61 → 86
27 → 90
Boolean
.50
97 → 97
99 → 99
Purchase
Menu
.03
94 → 86
24 → 79
T/F
.80
45 → 73
17 → 82
Boolean
.90
73 → 87
76 → 97
Table 2: Original → calibrated correct counts. Arithmetic and purchase denominators are 100 per coverage; time and counting denominators are 80. Lookup remains 100/100 throughout at τ=.5 . No evaluation labels are used to select thresholds.
Large language models can interpret natural lan- guage, yet robust decisions remain challenging. Jev-like models expose structured choices, but these interfaces do not directly provide numeri- cal values at a requested precision. We propose NUMERICJEV, a training-free numerical decod- ing algorithm that enables numerical output from any LLM with a Jev-like structured-choice in- terface. Surprisingly, on our arithmetic bench- mark, it outperforms direct selection from a can- didate list containing the correct answer by 2.93 percentage points (Figure 1). Our motivation comes from the observation that numerical range selection is itself a decision problem that Jev- like LLMs can address. NUMERICJEV recur- sively refines a range through a multiway deci- sion tree while retaining the original question in context, without parameter updates or hidden- state access. On a 100-value grid, a ten-way tree requires only two decision rounds. Range- normalized MAE is 1.84% versus 5.18% for di- rect choice. A separate three-date historical- index study yields 4.58% mean relative recall er- ror and 0% readout error when the value is sup- plied. Code is available at https://github. com/Bring-AI/jev-numeric.
LLM-as-a-judge scales evaluation, but reasoning judges are slow and costly. We study JEV-as-a-Judge: evaluation with JEV, a decision-only judge that returns label probabilities instead of text, and whose confidence decides whether to accept its verdict or escalate to a reasoning judge. Against sixteen generative and reward-model judges, with blinded human adjudication, JEV comes within three points of GPT-6 wherever a verdict can be read off the text, at 0.36% of its fee and a 0.15-second median latency, and falls behind where the verdict must be derived, as in math, code, and logic. Its confidence marks this boundary. With a threshold frozen in advance, accepting confident verdicts and escalating the rest is 0.9 points more accurate than GPT-6 on 1,610 held-out pairs at 41% of its fee, and in a pre-specified live test on two new workloads the cascade matches GPT-6's accuracy exactly. Confidence routing weakens on style-adversarial pairs and reference-free prose; we close with a simple recipe for validating thresholds locally.
Jev-style decision models return categorical probability distributions over predefined options without generating free-form text, enabling software systems to act on their outputs directly. In this work, we investigate the extent to which general-purpose LLMs already possess this capability out of the box, and when fine-tuning is actually necessary. We present LLM2Jev, an architecture-preserving framework that extracts calibrated decisions directly from next-token probabilities over bracketed numeric identifiers. LLM2Jev provides both a training-free inference recipe and a fine-tuning objective that optimizes candidate selection via a tree-factorized listwise loss while anchoring auxiliary predictions to the base model using KL divergence penalties. Evaluating on Qwen3.5-4B and Qwen3-0.6B, we find that modern LLMs are inherently effective decision models: without training, the 4B model matches community Jev-style models built on the same backbone, outperforms letter-logit readouts, supports arbitrary option counts, and natively handles multimodal decisions over images. Fine-tuning provides targeted rather than universal benefits -- substantially improving weaker models and specific tasks (such as many-option intent routing), but offering diminishing returns for strong backbones. Crucially, our KL anchors prevent behavioral degradation in conversational text generation, with LoRA delivering the strongest performance on capable models.