LLM2Jev: LLMs Are Already Jev-Style Decision Models -- When and How to Fine-Tune Them
Organizations: Microsoft
Abstract
Jev-style decision models return categorical probability distributions over predefined options without generating free-form text, enabling software systems to act on their outputs directly. In this work, we investigate the extent to which general-purpose LLMs already possess this capability out of the box, and when fine-tuning is actually necessary. We present LLM2Jev, an architecture-preserving framework that extracts calibrated decisions directly from next-token probabilities over bracketed numeric identifiers. LLM2Jev provides both a training-free inference recipe and a fine-tuning objective that optimizes candidate selection via a tree-factorized listwise loss while anchoring auxiliary predictions to the base model using KL divergence penalties. Evaluating on Qwen3.5-4B and Qwen3-0.6B, we find that modern LLMs are inherently effective decision models: without training, the 4B model matches community Jev-style models built on the same backbone, outperforms letter-logit readouts, supports arbitrary option counts, and natively handles multimodal decisions over images. Fine-tuning provides targeted rather than universal benefits -- substantially improving weaker models and specific tasks (such as many-option intent routing), but offering diminishing returns for strong backbones. Crucially, our KL anchors prevent behavioral degradation in conversational text generation, with LoRA delivering the strongest performance on capable models.
Figures & tables
| Config | Data | Examples | Steps |
| Intent-100 | intent set | 3,000 | 100 |
| Intent-1ep | intent set | 95,461 | 2,984 |
| Reason-9k | R6 + I3 (1k each) | 9,000 | 282 |
| Reason-21k | R6 (3k) + I3 (1k) | 21,000 | 657 |
| Reason-30k | Reason-21k + 3 reasoning sets (3k) | 30,000 | 938 |
| Reason+Long-12k | Reason-9k + 3 long-input sets | 12,000 | 375 |
| JevBench | External | ||||
| System | Acc. | Hard | ECE | Macro | B77 |
| GPT-5.6 Sol (generation) | 94.4 | 88.3 | – | 91.5 | 86.0 |
| Winnow-12B † | 86.6 | 74.8 | .058 | 82.2 | – |
| 4B training-free (ours) | 81.4 | 64.9 | .057 | 78.6 | 69.0 |
| 4B fine-tuned (ours) † | 80.5 | 61.3 | .079 | 78.3 | 75.9 |
| 4B LoRA (ours) † | 84.0 | 67.6 | .067 | 79.2 | 74.2 |
Appendix figures & tables14 assets
Supplementary material from the paper’s appendix.
Appendix
| Item | Setting |
|---|---|
| Trainable parameters | language model + LM head (all) |
| LoRA (§ 4.4.3 ) | rank 16, , dropout 0, on decoder linear layers (32.5M trainable parameters for 4B, 10.1M for 0.6B); merged after training |
| Parallelism, precision | FSDP; BF16 compute, FP32 optimizer states |
| Optimizer | AdamW; weight decay 0.01; gradient clipping 1.0 |
| Peak learning rate | 4B: ; 0.6B: ; LoRA: |
| Schedule | linear warmup (10% of steps; 100 for one-epoch runs), linear decay to 10% of the peak; one pass over the data |
| Dataset | Task | Options | Intent | Reason-9k | Reason-21k | Reason-30k | Reason+Long-12k | License |
| CLINC150 | intent (10 domains) | 151 | 15,207 | 1k | 1k | 1k | 1k | CC BY 3.0 |
| MASSIVE (en-US) | intent (voice assistant) | 60 | 11,390 | 1k | 1k | 1k | 1k | CC BY 4.0 |
| Bitext | intent (e-commerce support) | 27 | 21,839 | 1k | 1k | 1k | 1k | CDLA-Sharing-1.0 |
| CommonsenseQA | commonsense QA | 5 | 9,139 | – | – | – | – | MIT |
| HellaSwag | situation completion | 4 | 37,886 | – | – | – | – | MIT |
| ReClor | argument reasoning | 4 | – | 1k | 3k | 3k | 1k | research only |
| Suite | Dataset | License |
| JevBench | 231 public items | MIT |
| External | BoolQ | CC BY-SA 3.0 |
| MMLU | MIT | |
| MMLU-Pro | MIT | |
| ARC-Challenge | CC BY-SA 4.0 | |
| WinoGrande | CC BY |
| Intent | Reasoning | JevBench | |
|---|---|---|---|
| Examples | 95,461 | 9,000 | 231 |
| Intent classification | 51% | 33% | 10% |
| Two-option | 0% | 22% | 32% |
| Verification | 0% | 11% | 32% |
| Answer “none”/“unknown” | 0.3% | 6.0% | 3.5% |
| Training set | B77 | BoolQ | MMLU | Pro | ARC | Wino | SciQ | JevB |
|---|---|---|---|---|---|---|---|---|
| Intent | .553 | .243 | .241 | .256 | .227 | .098 | .189 | .295 |
| Reasoning | .444 | .332 | .337 | .322 | .264 | .132 | .222 | .258 |
| JevBench (231 public items) | External | General | |||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|
| Model | Config | Steps | Acc. | Won/lost ( ) | Hard | Long/37 | ECE | “Yes” | Macro | B77 | Avg. |
| 4B | original | 0 | 81.4 | – | 72 | 21 | .057 | 36 | 78.6 | 69.0 | 67.3 |
| Intent-100 | 100 | 79.2 | 13/18 (.47) | 69 | 22 | .106 | 54 | 78.3 | 74.7 | 68.3 | |
| Intent-1ep | 2,984 | 76.2 | 10/22 (.05) | 63 | 18 | .116 | 50 | 77.7 | 75.5 | 68.6 | |
| Reason-9k | 282 | 80.5 | 12/14 (.85) | 71 | 24 | .046 | 42 | 78.8 | 74.6 | 68.5 | |
| Reason-21k | 657 | 78.4 | 10/17 (.25) | 63 | 17 | .086 | 43 | 78.8 | 74.1 | 68.7 | |
| System / config | BoolQ | MMLU | Pro | ARC-C | Wino. | SciQ | Macro | B77 | ECE | |
|---|---|---|---|---|---|---|---|---|---|---|
| Reference | GPT-5.6 Sol (generation) | 91.80 | 91.42 | 77.62 | 98.17 | 90.40 | 99.50 | 91.48 | 86.04 | – |
| Winnow-12B (Q8) | 89.20 | 79.67 | 56.75 | 95.17 | 72.60 | 100.00 | 82.23 | – | .080 | |
| reflex 4B (LoRA) | 89.00 | 75.00 | 52.50 | 92.00 | 64.80 | 99.50 | 78.80 | – | .056 | |
| SemIf | 88.00 | 71.50 | 40.25 | 90.00 | 62.80 | 99.75 | 75.38 | – | .080 | |
| open-alternative-jev | 87.80 | 71.92 | 41.50 | 89.50 | 63.00 | 99.50 | 75.54 | – | .096 | |
| 4B | original (training-free) | 88.00 | 77.08 | 47.12 | 93.33 | 66.20 | 99.75 | 78.58 | 69.03 | .058 |
| Config | Steps | Avg. | GSM8K | IFEval | TriviaQA | LAMBADA | LAMBADA ppl | WikiText ppl | |
|---|---|---|---|---|---|---|---|---|---|
| 4B | original | 0 | 67.3 | 84.0 | 80.5 | 41.3 | 63.2 | 5.11 | 11.47 |
| Intent-100 | 100 | 68.3 | 84.4 | 83.0 | 42.3 | 63.6 | 5.08 | 11.44 | |
| Intent-1ep | 2,984 | 68.6 | 85.2 | 84.0 | 41.7 | 63.6 | 5.19 | 11.56 | |
| Reason-9k | 282 | 68.5 | 85.6 | 83.5 | 41.3 | 63.4 | 5.15 | 11.44 | |
| Reason-21k | 657 | 68.7 | 87.2 | 84.0 | 40.3 | 63.4 | 5.12 | 11.45 | |
| Reason-30k | 938 | 68.6 | 86.0 | 82.5 | 42.3 | 63.6 | 5.23 | 11.50 |
| MMBench-EN v1.1 dev (1,292 questions) | MMStar (1,498 questions) | |||||||
| Model | Acc. | Won/lost | Coarse | Same | ECE | Acc. | Won/lost | ECE |
| original | 82.4 | – | 81.2 | 87.6 | .014 | 64.6 | – | .090 |
| Intent-1ep | 83.6 | 46/30 | 85.4 ∗ | 89.2 | .011 | 64.2 | 110/115 | .073 |
| Reason-30k | 85.1 ∗ | 54/18 | 85.1 ∗ | 91.4 | .016 | 65.0 | 98/91 | .051 |
| Reason+Long-12k | 84.6 ∗ | 44/15 | 84.8 ∗ | 91.2 | .014 | 65.6 | 103/87 | .077 |
| original, no image | 15.2 | – | 9.1 | 30.8 | .214 | 28.8 | – | .234 |
| JevBench | External | General | Behavior | ||||||
|---|---|---|---|---|---|---|---|---|---|
| Model | Method | Acc. | ECE | Macro | B77 | Avg. | Text | Chat | |
| 4B | original | – | 81.4 | .057 | 78.6 | 69.0 | 67.3 | 0 | – |
| Full | 1 | 80.5 | .079 | 78.3 | 75.9 | 67.4 | 0 | 16/20 | |
| Full | 0.1 | 79.7 | .087 | 79.3 | 74.9 | 68.7 | 0 | 24/21 | |
| Full | 0.01 | 79.2 | .081 | 79.3 | 74.6 | 69.8 ∗ | 1 | 23/20 | |
| Full | 0 | 79.7 | .084 | 79.6 | 74.2 | 69.1 ∗ | 69 | 19/21 | |
| External | General | |||||||||||||||
| Model | Method | BoolQ | MMLU | Pro | ARC-C | Wino. | SciQ | Macro | B77 | GSM8K | IFEval | TriviaQA | LAMBADA | Avg. | WikiText ppl | |
| 4B | original | – | 88.0 | 77.1 | 47.1 | 93.3 | 66.2 | 99.8 | 78.6 | 69.0 | 84.0 | 80.5 | 41.3 | 63.2 | 67.3 | 11.47 |
| Full | 1 | 87.6 | 77.2 | 45.0 | 91.8 | 68.2 | 100.0 | 78.3 | 75.9 | 86.0 | 80.0 | 40.7 | 63.0 | 67.4 | 11.46 | |
| Full | 0.1 | 88.4 | 77.4 | 46.9 | 93.0 | 70.2 | 99.8 | 79.3 | 74.9 | 87.2 | 82.5 | 42.3 | 62.8 | 68.7 | 11.43 | |
| Full | 0.01 | 87.2 | 77.3 | 47.8 | 92.3 | 71.6 | 99.8 | 79.3 | 74.6 | 89.2 ∗ | 83.0 | 43.0 | 63.8 | 69.8 ∗ | 11.40 | |
| Full | 0 | 88.0 | 77.4 | 48.4 | 93.3 | 70.8 | 99.8 | 79.6 | 74.2 | 87.2 | 82.0 | 43.0 | 64.2 | 69.1 ∗ | 11.41 | |
| Drift (development sets) | JevBench text mode | Chat (560 replies) | ||||||||||
| Model | Method | KL start | KL answer | Tree loss | Parse fail. | At limit | No end | New/res. | Loops | New/res. | Shorter/longer | |
| 4B | original | – | 0 | 0 | 0.763 | 0 | 0 | 69 | – | 1 | – | – |
| Full | 1 | .0010 | .0081 | 0.424 | 0 | 0 | 65 | 16/20 | 1 | 1/1 | 106/33 ∗ | |
| Full | 0.1 | .0023 | .018 | 0.422 | 0 | 0 | 72 | 24/21 | 2 | 2/1 | 112/29 ∗ | |
| Full | 0.01 | .0043 | .033 | 0.412 | 1 | 0 | 72 | 23/20 | 2 | 2/1 | 116/26 ∗ | |
| Full | 0 | .048 | .305 | 0.421 | 69 | 44 | 67 | 19/21 | 2 | 2/1 | 105/36 ∗ | |
| System | Base | Training | Readout | Public | Sealed |
| Imajev-4B ( mohit67890, 2026 ) | Qwen3.5-4B | LoRA in four stages (human data, pseudo-labels, teacher questions, soft labels), weight-averaged | new head + temperature | 86.1 | 37.0 |
| Plumb-4B ( crh225, 2026 ) | JevK5 v0.2 | LoRA over five rounds: teacher questions, mined errors, long documents, replayed public data | letter logits + temperature | 89.6 | 38.0 |
| decider-4b v2 ( Mapika, 2026 ) | Qwen3.5-4B-Base | full SFT, then LoRA; about 95 public sets, programmatic and teacher data | letter logits + temperature | 83.5 | 34.7 |
| Jev 1.13.0 ( Almeida, 2026 ) | undisclosed | RLCD (details undisclosed) | – | 86.6 | 36.7 |
| JevK5 v0.2 ( allebee, 2026 ) | Qwen3.5-4B | distilled LoRA: 3,272 teacher questions plus equal public replay | letter logits + temperature | 85.3 | 33.1 |
| Cygnet ( blockbrain-ai, 2026 ) | Gemma-4-12B-it | none | letter logits + temperature | 87.9 | 33.8 |