JevOut: Natural Context Can Flip Decision Models
Organizations: University of Southern California
Abstract
Dedicated decision models such as Jev map unstructured language to probability distributions over finite choices, allowing their outputs to directly route requests, select tools, and trigger actions. Yet real-world inputs rarely arrive in isolation: they come with background details and surrounding context. We find that short additions that fit naturally into this context can nevertheless redirect an otherwise correct decision, even when the correct answer remains unchanged. To study this behavior, we fix a wrong target option for each initially correct item and use the model's option probabilities to refine fluent context additions while preserving the source, question, choices, and gold answer. Within 64 accepted target evaluations, the optimizer identifies contexts that redirect Jev on 312 of 508 initially correct decisions (61.4%); in 229 cases, Jev assigns at least 0.7 probability to the fixed wrong option. Across seven datasets, three additional decision systems show targeted flip rates of 64.9%-73.2% on decisions they initially answer correctly. Taken together, these results expose a pronounced fragility in current decision models: short, ordinary-looking context can shift a correct choice to a high-confidence wrong one. Because these models turn language directly into downstream choices, this sensitivity raises concerns about treating their probability outputs as reliable decision interfaces.
Figures & tables
| Target | One-shot | Context optimization | |||
| Neutral | Target-aware | TFR | |||
| Jev | 508 | 2.2 (11) | 16.9 (86) | 61.4 (312) | 45.1 (229) |
| OpenSourceJev | 328 | 6.4 (21) | 16.2 (53) | 72.6 (238) | 64.3 (211) |
| Von | 328 | 7.6 (25) | 21.6 (71) | 73.2 (240) | 54.0 (177) |
| Plain Qwen | 285 | 8.4 (24) | 18.6 (53) | 64.9 (185) | 60.4 (172) |
| Proposer | One generation | Best-of-four generations | |
| TFR | TFR | ||
| Base | 16.1 (82) | 29.1 (148) | 17.9 (91) |
| V1: transition-trained | 18.7 (95) | 32.9 (167) | 22.0 (112) |
| V2: context-trained | 21.9 (111) | 34.3 (174) | 21.3 (108) |
Appendix figures & tables21 assets
Supplementary material from the paper’s appendix.
Appendix
| Dataset | Jev | OpenSourceJev | Von | plain Qwen |
| MMLU-Pro | 83.0 / 83 | 33.0 / 33 | 18.0 / 18 | 28.0 / 28 |
| SuperGPQA | 46.0 / 46 | 12.0 / 12 | 16.0 / 16 | 16.0 / 16 |
| MuSR | 61.0 / 61 | 36.0 / 36 | 60.0 / 60 | 46.0 / 46 |
| ToMBench | 72.0 / 72 | 37.0 / 37 | 41.0 / 41 | 40.0 / 40 |
| LAR-ECHR | 80.0 / 80 | 46.0 / 46 | 47.0 / 47 | 51.0 / 51 |
| SATA | 82.3 [42.0] / 96 | 78.2 [25.0] / 88 | 75.3 [28.0] / 89 | 74.5 [19.0] / 82 |
| Target | Backbone / service | Finite distribution | Role in the study |
| Jev | hosted jev-1.13.0 | Native Choice probabilities or a binary Noul probability | Primary decision-model target |
| OpenSourceJev | Qwen3-1.7B Q8 | Conditional branch-token likelihoods with released Noul calibration | Open Jev-style implementation |
| Von | wfzyx/von-1.0 , 395M | Native non-autoregressive typed distribution | Architecturally distinct decision model |
| plain Qwen | Qwen3-1.7B Q8 | Conditional likelihood of numbered option labels | Matched-backbone language-model scorer |
| Experiment | Population | Generations | Target evaluations |
| Primary optimization | 1,449 | 68,912 | 60,685 |
| Formal one-shot controls | 1,449 | 2,898 | 2,778 |
| Feedback ablations | 140 | 14,722 | 13,197 |
| Directed transfer | 12 pairs | — | 2,657 |
| Development construction | 260 | 13,989 | 12,239 |
| Single-addition Base/V1 | 508 | 4,064 | 3,787 |
| Dataset | Neutral | Target-aware | TFR (count) | 95% CI | ||
| MMLU-Pro | 83 | 2.4 | 12.0 | 45.8 (38) | [35.5, 56.4] | 31.3 |
| SuperGPQA | 46 | 8.7 | 15.2 | 67.4 (31) | [53.0, 79.1] | 34.8 |
| MuSR | 61 | 3.3 | 11.5 | 62.3 (38) | [49.7, 73.4] | 42.6 |
| ToMBench | 72 | 2.8 | 11.1 | 76.4 (55) | [65.4, 84.7] | 63.9 |
| LAR-ECHR | 80 | 0.0 | 3.8 | 40.0 (32) | [30.0, 51.0] | 21.2 |
| SATA | 96 | 0.0 | 24.0 | 60.4 (58) | [50.4, 69.6] | 49.0 |
| Dataset | Neutral | Target-aware | TFR (count) | 95% CI | ||
| MMLU-Pro | 33 | 6.1 | 12.1 | 69.7 (23) | [52.7, 82.6] | 69.7 |
| SuperGPQA | 12 | 16.7 | 16.7 | 75.0 (9) | [46.8, 91.1] | 66.7 |
| MuSR | 36 | 8.3 | 5.6 | 61.1 (22) | [44.9, 75.2] | 61.1 |
| ToMBench | 37 | 8.1 | 16.2 | 75.7 (28) | [59.9, 86.6] | 73.0 |
| LAR-ECHR | 46 | 4.3 | 6.5 | 39.1 (18) | [26.4, 53.5] | 39.1 |
| SATA | 88 | 6.8 | 25.0 | 80.7 (71) | [71.2, 87.6] | 53.4 |
| Dataset | Neutral | Target-aware | TFR (count) | 95% CI | ||
| MMLU-Pro | 18 | 11.1 | 44.4 | 88.9 (16) | [67.2, 96.9] | 72.2 |
| SuperGPQA | 16 | 12.5 | 50.0 | 100.0 (16) | [80.6, 100.0] | 50.0 |
| MuSR | 60 | 8.3 | 15.0 | 55.0 (33) | [42.5, 66.9] | 36.7 |
| ToMBench | 41 | 2.4 | 24.4 | 78.0 (32) | [63.3, 88.0] | 63.4 |
| LAR-ECHR | 47 | 8.5 | 8.5 | 61.7 (29) | [47.4, 74.2] | 51.1 |
| SATA | 89 | 6.7 | 14.6 | 76.4 (68) | [66.6, 84.0] | 68.5 |
| Dataset | Neutral | Target-aware | TFR (count) | 95% CI | ||
| MMLU-Pro | 28 | 10.7 | 21.4 | 57.1 (16) | [39.1, 73.5] | 50.0 |
| SuperGPQA | 16 | 25.0 | 25.0 | 56.2 (9) | [33.2, 76.9] | 50.0 |
| MuSR | 46 | 10.9 | 10.9 | 65.2 (30) | [50.8, 77.3] | 58.7 |
| ToMBench | 40 | 0.0 | 12.5 | 60.0 (24) | [44.6, 73.7] | 57.5 |
| LAR-ECHR | 51 | 3.9 | 9.8 | 43.1 (22) | [30.5, 56.7] | 35.3 |
| SATA | 82 | 12.2 | 30.5 | 89.0 (73) | [80.4, 94.1] | 86.6 |
| Paired system | Matched | Jev TFR (%) | Other TFR (%) |
| OpenSourceJev | 277 | 58.1 | 70.8 |
| Von | 268 | 58.6 | 73.5 |
| Plain Qwen | 243 | 53.5 | 64.6 |
| Target | 16 calls | 32 calls | 48 calls | 64 calls |
| Jev | 40.4 | 51.4 | 57.3 | 61.4 |
| OpenSourceJev | 50.9 | 61.0 | 68.3 | 72.6 |
| Von | 52.4 | 64.9 | 70.7 | 73.2 |
| Plain Qwen | 42.5 | 55.8 | 60.7 | 64.9 |
| Calls | Full | Probability-only | Label-only | Difference | 95% CI |
| 16 | 37.9 | 38.6 | 37.9 | 0.7 | [0.0, 2.1] |
| 32 | 50.7 | 50.0 | 47.9 | 2.1 | [0.0, 5.0] |
| 48 | 57.1 | 55.7 | 52.1 | 3.6 | [0.0, 7.9] |
| 64 | 61.4 | 63.6 | 56.4 | 7.1 | [2.1, 12.9] |
| Source | Destination | Flips | Rate (%) | 95% CI | |
| Jev | OpenSourceJev | 240 | 75 | 31.2 | [25.7, 37.4] |
| Jev | Von | 256 | 65 | 25.4 | [20.4, 31.1] |
| Jev | Plain Qwen | 219 | 57 | 26.0 | [20.7, 32.2] |
| OpenSourceJev | Jev | 274 | 69 | 25.2 | [20.4, 30.6] |
| OpenSourceJev | Von | 202 | 58 | 28.7 | [22.9, 35.3] |
| OpenSourceJev | Plain Qwen | 224 | 99 | 44.2 | [37.8, 50.7] |
| Dataset | OpenSourceJev | Von | Plain Qwen |
| MMLU-Pro | 6/29 (20.7) | 8/17 (47.1) | 5/26 (19.2) |
| SuperGPQA | 1/5 (20.0) | 3/8 (37.5) | 1/10 (10.0) |
| MuSR | 5/28 (17.9) | 4/39 (10.3) | 5/31 (16.1) |
| ToMBench | 4/33 (12.1) | 12/34 (35.3) | 7/35 (20.0) |
| LAR-ECHR | 6/42 (14.3) | 6/40 (15.0) | 4/46 (8.7) |
| SATA | 33/51 (64.7) | 22/76 (28.9) | 32/58 (55.2) |
| Dataset | Jev | Von | Plain Qwen |
| MMLU-Pro | 0/29 (0.0) | 5/10 (50.0) | 11/24 (45.8) |
| SuperGPQA | 1/5 (20.0) | 2/2 (100.0) | 3/6 (50.0) |
| MuSR | 7/28 (25.0) | 3/22 (13.6) | 9/28 (32.1) |
| ToMBench | 9/33 (27.3) | 11/25 (44.0) | 13/30 (43.3) |
| LAR-ECHR | 5/42 (11.9) | 1/25 (4.0) | 7/42 (16.7) |
| SATA | 25/85 (29.4) | 26/72 (36.1) | 50/75 (66.7) |
| Dataset | Jev | OpenSourceJev | Plain Qwen |
| MMLU-Pro | 1/17 (5.9) | 3/10 (30.0) | 0/6 (0.0) |
| SuperGPQA | 3/8 (37.5) | 0/2 (0.0) | 1/3 (33.3) |
| MuSR | 6/39 (15.4) | 4/22 (18.2) | 4/26 (15.4) |
| ToMBench | 8/34 (23.5) | 3/25 (12.0) | 4/26 (15.4) |
| LAR-ECHR | 2/40 (5.0) | 1/25 (4.0) | 1/28 (3.6) |
| SATA | 27/80 (33.8) | 29/56 (51.8) | 26/68 (38.2) |
| Dataset | Jev | OpenSourceJev | Von |
| MMLU-Pro | 2/26 (7.7) | 9/24 (37.5) | 3/6 (50.0) |
| SuperGPQA | 2/10 (20.0) | 3/6 (50.0) | 3/3 (100.0) |
| MuSR | 11/31 (35.5) | 12/28 (42.9) | 2/26 (7.7) |
| ToMBench | 7/35 (20.0) | 13/30 (43.3) | 7/26 (26.9) |
| LAR-ECHR | 7/46 (15.2) | 9/42 (21.4) | 3/28 (10.7) |
| SATA | 25/78 (32.1) | 41/58 (70.7) | 24/76 (31.6) |
| Dataset | One proposal | Best-of-four | |||
| Base | V1 | Base | V1 | ||
| MMLU-Pro | 83 | 12.0% | 10.8% | 18.1% | 20.5% |
| SuperGPQA | 46 | 10.9% | 26.1% | 32.6% | 39.1% |
| MuSR | 61 | 11.5% | 19.7% | 21.3% | 32.8% |
| ToMBench | 72 | 11.1% | 13.9% | 18.1% | 25.0% |
| LAR-ECHR | 80 | 0.0% | 8.8% | 12.5% | 20.0% |
| Dataset | Base | V1 | V2 | |
| MMLU-Pro | 83 | 10.8 (9) | 13.3 (11) | 13.3 (11) |
| SuperGPQA | 46 | 10.9 (5) | 19.6 (9) | 8.7 (4) |
| MuSR | 61 | 9.8 (6) | 18.0 (11) | 23.0 (14) |
| ToMBench | 72 | 9.7 (7) | 15.3 (11) | 16.7 (12) |
| LAR-ECHR | 80 | 3.8 (3) | 6.2 (5) | 13.8 (11) |
| SATA | 96 | 22.9 (22) | 18.8 (18) | 27.1 (26) |
| Dataset | Base | V1 | V2 | |
| MMLU-Pro | 83 | 21.7 (18) | 25.3 (21) | 24.1 (20) |
| SuperGPQA | 46 | 28.3 (13) | 30.4 (14) | 39.1 (18) |
| MuSR | 61 | 29.5 (18) | 39.3 (24) | 32.8 (20) |
| ToMBench | 72 | 23.6 (17) | 33.3 (24) | 34.7 (25) |
| LAR-ECHR | 80 | 15.0 (12) | 16.2 (13) | 25.0 (20) |
| SATA | 96 | 32.3 (31) | 34.4 (33) | 34.4 (33) |
| Comparison | Generations | Difference | 95% CI | |
| V2 minus Base | One | 5.7 | [2.4, 9.1] | 0.0013 |
| V2 minus Base | Four | 5.1 | [1.8, 8.5] | 0.0038 |
| V2 minus V1 | One | 3.1 | [0.0, 6.3] | 0.0764 |
| V2 minus V1 | Four | 1.4 | [-2.0, 4.7] | 0.4944 |
| Proposer | Generated | Accepted | First accepted | Mean additions | Four: |
| Base | 2032 | 1882 | 480 | 1.64 | 17.9 (91) |
| V1 | 2032 | 1875 | 467 | 1.78 | 22.0 (112) |
| V2 | 2032 | 1808 | 453 | 1.76 | 21.3 (108) |
| Target | Contexts | 1 | 2 | 3 | 4 | Median words | Word IQR |
| Jev | 312 | 162 | 91 | 47 | 12 | 31 | 21–54 |
| OpenSourceJev | 238 | 148 | 47 | 28 | 15 | 28 | 21–44 |
| Von | 240 | 144 | 68 | 23 | 5 | 28 | 22–45 |
| Plain Qwen | 185 | 113 | 46 | 17 | 9 | 28 | 22–45 |