Large language models (LLMs) annotate and scale political text or constructs by generating text tokens. A new class of models, which TypeSafe markets as "System One" models, instead returns decisions and probability distributions across a user-supplied fixed answer set. A commercial model, JEV, is advertised as having a dramatic cost and speed advantage over traditional LLMs along with better calibrated decisions. As such, it might be useful for social scientists looking to quickly and cost-effectively annotate or scale large corpora of text and have a reliable indicator of a classifier's uncertainty. Yet, the accuracy of these claims and the broader model accuracy in social science text-based tasks are not yet established. In this paper, we do just that and hope to establish the suitability of JEV for social science tasks. We compare JEV with LLMs and human coders from published research, and with a current mid-tier commercial LLM (GPT-6 Luna) and an open-weight alternative (Qwen3.8-27B). We find that JEV matches, or comes close to, the capabilities of both LLMs in a variety of tasks. However, we find no cost advantage over GPT-6 Luna at OpenAI's batch prices. Further, we find that, when each question is asked once, JEV's probabilities are better calibrated than GPT-6 Luna's token probabilities, but not consistently better than Qwen3.8-27B's. We conclude that unless researchers have a need for speed, JEV's only obvious advantage is ease of parsing the underlying choice probabilities.
Figures & tables
JEV compared with
Study
Data
Test replicated
Human reference
published models
GPT-6 Luna
Qwen 27B
Annotation
Gilardi et al. (2023)
Tweets on content moderation, 2021 and 2023
Relevance (relevant or not), four runs; accuracy
Trained annotators; Mechanical Turk workers as a comparator
Better than ChatGPT
Worse
Slightly better (2021); worse (2023) a
Ornstein et al. (2025)
945 tweets on two Supreme Court decisions
Sentiment (positive, neutral, negative); correlation of the probability-based score
Mean of three expert ratings
Worse than GPT-3 and GPT-4; better than Twitter RoBERTa and Naive Bayes
Worse
Worse
Brandt et al. (2026)
37,709 Global Terrorism Database incidents; 322 BBC news articles
Attack type, nine categories, names only; accuracy and macro F1. BBC news: conflict F1
GTD coders’ attack type
Better than prompted LLMs; worse than fine-tuned ConfliBERT and ConflLlama. BBC: matches Llama 3.1, worse than ConfliBERT
Better acc.; slightly better F1. BBC: better
Better acc. and F1. BBC: better
Weidmann et al. (2026)
53 V-Dem indicators, 171 countries, 2023 (recall)
Ordinal indicator coded from the country name; correlation and exact agreement
V-Dem expert codes
Better than GPT-4o and Llama-3.1 70B
Better
Better
Table 1: The seven published applications re-run with JEV. The human reference is what every model is scored against; designs are unchanged except as described in the text. Reading tasks give the model a text, recall tasks only a name. The last three columns compare JEV with the published models on each study’s headline measure, and with GPT-6 Luna and the open-weight Qwen3.8-27B (reasoning off) run under the same prompt (paired bootstrap, Efron 1979 ; details in Tables 9 and 10 ). Against GPT-6 Luna and Qwen, better or worse means the paired 95% interval excludes zero; slightly better or worse means JEV’s estimate is higher or lower but the interval includes zero. Against the published models, which have no paired intervals, better or worse means above or below their reported figures, and within range means between the lowest and highest. a The interval’s lower bound is 0.2 points and depends on the bootstrap seed. BBC conflict F1: JEV 0.321, Llama 3.1 0.322, ConfliBERT 0.681.
$ per 1,000 decisions
Seconds per decision
Decisions per second
Application
JEV
Luna (Batch)
Luna (Std.)
Qwen
JEV
Luna
Qwen
JEV
Luna
Qwen
Gilardi et al. (2023)
0.017
0.017
0.034
0.085
0.09
0.77
0.44
156
18
31
Ornstein et al. (2025)
0.024
0.012
0.025
0.128
0.26
0.75
0.39
47
18
30
Brandt et al. (2026)
0.027
0.018
0.035
0.139
0.27
1.04
0.87
51
14
14
Weidmann et al. (2026) a
0.020
0.012
0.024
0.058
0.26
0.75
0.41
45
17
31
Di Leo et al. (2025)
0.018
0.008
0.016
0.042
0.25
0.78
0.43
46
20
28
Table 2: Cost and speed of JEV, GPT-6 Luna and Qwen3.8-27B on the primary task of each application. A decision is one answer to one question about one item. Cost: dollars per 1,000 decisions; JEV and Qwen as billed, GPT-6 Luna at OpenAI’s posted Batch and Standard rates. Qwen ran through OpenRouter on one host (Parasail, FP8) with no batch rate, so its cost is that of this route, not of running the open weights oneself. Speed: median seconds per decision with one request at a time, and decisions per second with 16 at a time, from 200 timed items per application sent from the Netherlands on 29–30 September 2026. Throughput depends on the route and rate limits and is not a property of the models. a Costs for JEV and Luna are from the run’s totals.
Figure 1: Distribution of each model’s probability that the first party is the more right-wing, across all party comparisons in our replication of Di Leo et al. (2025) ( N = 347,801; one Qwen reply was lost). For JEV the probability comes from a two-option Choice; for GPT-6 Luna and Qwen3.8-27B, from their first-token probabilities.
Figure 2: Calibration of JEV, GPT-6 Luna and Qwen3.8-27B against human coders, by task, with 95% bootstrap intervals. Left: calibration error (ECE), the gap between the probability a model gives its answer and the share of such answers that agree with the human label, over ten probability bins; lower is better. Right: AUROC, the chance that a correct answer gets a higher probability than an incorrect one; 0.5 is chance. Top rows: each question asked once. Middle rows: probabilities averaged over the two presentation orders. Bottom rows: on attack type GPT-6 Luna and Qwen write their probabilities in the reply; party positions and democracy indicators are recall tasks. In the attack-type and party-position rows, which have tens or hundreds of thousands of items, the intervals are narrower than the markers. Appendix H describes the Licht and MARPOR tasks.
Appendix figures & tables12 assets
Supplementary material from the paper’s appendix.
Appendix
Probability of choice
Pairs
Flip rate
Position bias
<.50
5,537
0.3655
− 0.1104
.50–.70
2,984
0.1146
− 0.0721
.70–.85
2,881
0.0139
− 0.0263
.85–.95
3,862
0.0003
− 0.0079
>.95
12,103
0.0000
− 0.0005
Appendix
Table 3: Stability and position bias of JEV’s pairwise judgments in Study 1, by the probability JEV gives its chosen option (27,367 pairs asked in both orders with no tie). “Flip rate” is the share of pairs whose winner changes when the two statements are swapped. “Position bias” is the pull toward the first slot on the probability scale; a negative value favours the second.
Annotator
Accuracy
Intercoder agreement
GPT-6 Luna, single run
91.1
—
JEV 1.13, single run
89.0
—
Qwen3.8-27B, single run
88.2
—
JEV 1.13, four runs, unanimity required
88.2
98.7
ChatGPT (temp 1)
72.8
92.0
Mechanical Turk
71.2
79.2
Appendix
Table 4: Tweet relevance classification, 2021 sample (2,403 tweets with agreed gold labels). Rows for ChatGPT, Mechanical Turk and trained annotators are published figures from Gilardi et al. (2023) ; the others are ours. JEV’s intercoder agreement is agreement among its four runs. GPT-6 Luna and the open-weight Qwen3.8-27B were each run once with the published instruction (Appendix H).
Measure
Kind
Scoring rule
ρ with experts
Qwen3.8-27B (few-shot)
generative
positive − negative
0.841
GPT-6 Luna (few-shot)
generative
positive − negative
0.816
GPT-3 (few-shot)
generative
principal component
0.807
GPT-6 Luna (zero-shot)
generative
positive − negative
0.791
GPT-4 (few-shot)
generative
positive − negative
0.789
JEV (few-shot)
constrained output
positive − negative
0.740
Appendix
Table 5: Twitter sentiment, replicating Application 1 of Ornstein et al. (2025) : correlation with the mean of three expert ratings across the 907 tweets on which every measure is available. Their four measures are recomputed from their files with their code and reproduce their published figures. Each is scored by the rule their figure code applies: the first principal component of the three class probabilities for GPT-3 and TweetNLP (the rule they pre-registered), positive minus negative for GPT-4. JEV is shown under both rules; GPT-6 Luna and the open-weight Qwen3.8-27B under GPT-4’s (Appendix H). Without the examples Qwen answered 941 of 945 tweets in prose, so its zero-shot score is not comparable.
Scorer
Mode
Expert pairs
Other models’ majority
Llama 3.1 405B
pairwise
0.815
0.949
Gemma 3 27B
pairwise
0.813
0.928
JEV
pairwise
0.796
0.910
GPT-4o
pairwise
0.817
0.906
GPT-4o mini
pairwise
0.824
0.899
Gemma 3 4B
pairwise
0.739
0.875
Appendix
Table 6: A/B correctness rather than correlation. Column 3 is accuracy on the expert-judged pairs; column 4 is agreement with the majority of the other six scorers across all 28,040 published pairs. JEV’s accuracy in column 3 is averaged over both presentation orders (0.800 in the published order alone; Table 8 ). Nine further models from an unpublished extension by the same authors drew their own pairs and share only 782 of those pairs, so they are not shown.
Attack type
Support
Confli- BERT
ConflLlama 8-bit
JEV names
JEV codebook
Qwen 3.5 names
Qwen 3.5 codebook
GPT-6 Luna names
GPT-6 Luna codebook
Qwen 27B names
Qwen 27B codebook
Assassination
2,990
0.74
0.59
0.29
0.54
0.02
0.38
0.48
0.66
0.28
0.75
Armed Assault
9,079
0.81
0.75
0.90
0.94
0.94
0.89
0.79
0.89
0.88
0.91
Bombing/Explosion
14,508
0.97
0.92
0.91
0.81
0.84
0.65
0.92
0.86
0.85
0.73
Hijacking
154
0.64
0.54
0.60
0.70
0.44
0.58
0.38
0.50
0.55
0.56
Hostage taking, barricade
230
0.38
0.27
0.33
0.46
0.26
0.37
0.23
0.46
0.29
0.41
Hostage taking, kidnapping
3,495
0.88
0.86
0.75
0.82
0.77
0.81
0.71
0.81
0.69
0.71
Appendix
Table 7: Recall by attack type on the 37,709 test incidents. Support is the number of incidents of that type. ConfliBERT and ConflLlama are recomputed from Brandt et al.’s per-incident predictions. GPT-6 Luna is a current generative model (Appendix H). Qwen 3.5 is Qwen3.5-9B and Qwen 27B is the open-weight Qwen3.8-27B. “Names” columns give the model only the category names; “codebook” columns add the GTD codebook definitions.
Measure
N
Precision
Recall
F1
95% CI
Accuracy
GPT-4o mini (paper)
460
0.835
0.845
0.840
[0.797, 0.876]
0.824
Gemma 3 27B (paper)
460
0.821
0.841
0.831
[0.787, 0.869]
0.813
GPT-4o (paper)
460
0.852
0.805
0.828
[0.780, 0.869]
0.817
JEV (Noul)
460
0.812
0.829
0.821
[0.772, 0.858]
0.802
Llama 3.1 405B (paper)
459
0.877
0.769
0.820
[0.770, 0.859]
0.815
JEV (rev)
460
0.799
0.841
0.819
[0.773, 0.858]
0.798
Appendix
Table 8: Agreement with expert raters: 460 judgments of 243 pairs. F1, defined as in the published article, treats “human picked statement 1” as the positive class, so it depends on order; accuracy is also shown. “fwd” is the published presentation order, “rev” the reverse and “(paper)” marks the published scorers. Intervals: 2,000 bootstrap resamples of pairs; rater clustering is not modelled. GPT-6 Luna and the open-weight Qwen3.8-27B are ours, in the published order (Appendix H).
Test
Statistic
GPT-6 Luna
JEV
Difference
Reading
Tweet relevance, 2021
accuracy (%)
91.1
89.0
2.1 [0.7, 3.5]
above
Tweet relevance, 2023
accuracy (%)
84.2
76.2
8.0 [4.1, 11.9]
above
Tweet sentiment, few-shot
ρ with experts
0.816
0.740
0.076 [0.039, 0.112]
above
GTD attack type, names
accuracy
0.709
0.715
− 0.006 [ − 0.009, − 0.003]
below
GTD attack type, names
macro F1
0.544
0.555
− 0.011 [ − 0.025, 0.002]
level
BBC conflict
conflict F1
0.139
0.321
− 0.182 [ − 0.308, − 0.050]
below
Appendix
Table 9: GPT-6 Luna against JEV on every test, under each paper’s published prompt (one pass, temperature 0, reasoning off; Appendix H). Differences are GPT-6 Luna minus JEV’s primary specification, with paired bootstrap 95% intervals; above, level and below say whether the interval lies above zero, contains it or lies below it. For calibration error lower is better, so “above” there favours JEV. The last three rows (democracy indicators) were added later.
Test
Statistic
Qwen
JEV
Difference
Reading
Tweet relevance, 2021
accuracy (%)
88.2
89.0
− 0.8 [ − 2.1, 0.4]
level
Tweet relevance, 2023
accuracy (%)
79.6
76.2
3.4 [0.2, 6.8]
above a
Problem frame, 2021
accuracy (%)
92.4
92.4
0.0 [ − 1.2, 1.2]
level
Problem frame, 2023
accuracy (%)
69.3
61.4
7.8 [1.2, 14.5]
above
Solution frame, 2021
accuracy (%)
81.4
84.7
− 3.4 [ − 4.4, − 2.3]
below
Solution frame, 2023
accuracy (%)
96.5
98.2
− 1.8 [ − 4.1, 0.0]
level
Appendix
Table 10: Qwen3.8-27B against JEV on every test, through OpenRouter on one host (Parasail, FP8), reasoning off, temperature 0, one pass, with the messages GPT-6 Luna received (Appendix H). Differences are Qwen3.8-27B minus JEV’s primary specification, with paired bootstrap 95% intervals; above, level and below say whether the interval lies above zero, contains it or lies below it. For calibration error lower is better. The frame rows use our reconstructed instructions. The survey calibration row averages over presentation orders and uses the top twenty reply tokens; the survey scale row is the correlation with GPT-4o’s Study 1 scale.
Macro F1
Unknown
Model
Accuracy
Nine types
Known types
recall
Fine-tuned on GTD labels (published)
ConfliBERT
0.840
0.744
0.775
0.550
ConflLlama, 8-bit
0.765
0.647
0.687
0.387
ConflLlama, 4-bit
0.729
0.571
0.608
0.323
Prompted with the codebook
Appendix
Table 11: Attack type of 37,709 GTD incidents, replicating Brandt et al. (2026) . Known types: macro F1 over the eight types, on the 33,381 incidents not coded Unknown. Unknown recall: share of the 4,328 Unknown incidents labelled Unknown. GPT-6 Luna was sent Qwen3.5-9B’s prompts, and the open-weight Qwen3.8-27B received GPT-6 Luna’s requests (Appendix H).
Coder
Correlation
Difference
Exact
Changed
Within one
Missing
JEV, most probable level
0.68 [0.66, 0.70]
− 0.16
54
36
92
0
JEV, expected level
0.75 [0.73, 0.76]
− 0.20
0
GPT-6 Luna
0.63 [0.61, 0.64]
− 0.09
52
38
91
0
Qwen3.8-27B
0.59 [0.57, 0.61]
− 0.17
51
35
89
0
Qwen3.8-27B, expected level
0.67 [0.66, 0.69]
− 0.18
0
GPT-4o
0.63 [0.62, 0.65]
− 0.28
49
36
90
5
Appendix
Table 12: V-Dem indicators coded from the country name alone, replicating Weidmann et al. (2026) : 9,041 country-indicator pairs (53 ordinal indicators, 171 countries, 2023), scored against V-Dem v14. Correlation is the mean over countries of the correlation across a country’s indicators, with a 95% country-bootstrap interval. Difference is the mean of the model’s code minus V-Dem’s (below zero is pessimistic). Exact, Changed (the 877 pairs whose value changed from 2022) and Within one are percentages of pairs; Missing counts pairs without a code. GPT-4o and Llama-3.1 rows rescore the published codes (the article reports 0.64 for GPT-4o). GPT-6 Luna and Qwen3.8-27B received the published prompt (Appendix H).
CHES experts
TEV voters
CMP manifestos
Year
GPT-3.5
JEV
GPT-6 Luna
GPT-3.5
JEV
GPT-6 Luna
GPT-3.5
JEV
GPT-6 Luna
1979
—
—
—
0.75
0.92 [0.79, 0.97]
0.93
0.49
0.57 [0.44, 0.69]
0.51
1984
0.77
0.88 [0.81, 0.92]
0.83
0.78
0.92 [0.84, 0.96]
0.89
0.53
0.59 [0.46, 0.69]
0.56
1989
0.77
0.88 [0.81, 0.92]
0.84
0.74
0.82 [0.65, 0.91]
0.79
0.60
0.59 [0.46, 0.69]
0.59
1994
0.81
0.89 [0.83, 0.93]
0.84
0.84
0.89 [0.79, 0.94]
0.83
0.44
0.52 [0.43, 0.60]
0.50
1999
0.75
0.88 [0.83, 0.91]
0.84
0.77
0.88 [0.81, 0.93]
0.85
0.44
0.49 [0.39, 0.57]
0.48
Appendix
Table 13: European party positions, replicating Di Leo et al. (2025) : Pearson correlation of Bradley-Terry scores with each benchmark by year. CHES: Chapel Hill Expert Survey (before 1999, the Ray–Marks–Steenbergen values); TEV: True European Voter surveys; CMP: Manifesto Project scores. Their GPT-3.5 figures, the modal answer over seven runs, are transcribed from their Figure 1; dashes mark panels they leave empty. JEV’s are a single pass with 95% Fisher intervals; GPT-6 Luna and, in panel b, Qwen3.8-27B are single passes without intervals (Appendix H). Panel b adds voters’ placements in the Comparative Study of Electoral Systems (CSES), available in their files for 1999 to 2014.
UK manifestos (experts)
EU speeches (crowd)
Economic
Social
Subsidies
JEV
0.97
0.92
0.81
JEV, relevance-weighted
0.97
0.92
0.86
GPT-4o (published)
0.98
0.94
0.93
GPT-6 Luna
0.98
0.95
0.95
Qwen3.8-27B
0.97
0.94
0.90
Appendix
Table 14: Party manifestos and legislative speeches, replicating Le Mens and Gallego (2025): Pearson correlation of document positions with the benchmark. Positions average the sentence scores of the sentences judged relevant; for JEV, those with a relevance probability of at least 0.5, and in the weighted row every sentence, weighted by that probability. Published figures are recomputed from their archive. Gaps carry 95% intervals from a paired bootstrap over documents. GPT-6 Luna and the open-weight Qwen3.8-27B received the published prompts, scale and NA option, and average the sentences they did not mark NA; their gaps with JEV use the 35 speeches on which JEV judged at least one sentence relevant.
Jev-style decision models return categorical probability distributions over predefined options without generating free-form text, enabling software systems to act on their outputs directly. In this work, we investigate the extent to which general-purpose LLMs already possess this capability out of the box, and when fine-tuning is actually necessary. We present LLM-as-Jev, an architecture-preserving framework that extracts calibrated decisions directly from next-token probabilities over bracketed numeric identifiers. LLM-as-Jev provides both a training-free inference recipe and a fine-tuning objective that optimizes candidate selection via a tree-factorized listwise loss while anchoring auxiliary predictions to the base model using KL divergence penalties. Evaluating on Qwen3.5-4B and Qwen3-0.6B, we find that modern LLMs are inherently effective decision models: without training, the 4B model matches community Jev-style models built on the same backbone, outperforms letter-logit readouts, supports arbitrary option counts, and natively handles multimodal decisions over images. Fine-tuning provides targeted rather than universal benefits -- substantially improving weaker models and specific tasks (such as many-option intent routing), but offering diminishing returns for strong backbones. Crucially, our KL anchors prevent behavioral degradation in conversational text generation, with LoRA delivering the strongest performance on capable models.
We ask whether Jev, a typed classifier that returns probabilities over permitted answers without generating text, can replace an LLM rubric judge. We compare it with three flash-tier LLM judges on nine panels drawn from seven benchmarks, giving every judge identical criterion texts. Jev's accuracy differs significantly from an LLM judge's in only 8 of 27 paired comparisons, ahead mostly on binary criteria and behind only on graded ones, and most of the other comparisons are inconclusive. Summed over the nine panels, the LLM judges, called once per criterion, cost 29 to 325 times as much as Jev and took 30 to 220 times as long. On graded criteria all four judges agree more with one another than with the labels and mostly assign lower levels than the raters. One of several observational accounts is that raters followed scale conventions our criterion texts omit. Jev's confidence ranks its own errors on most panels, which should make a cheap classifier the ideal first stage of a cascade that defers its uncertain verdicts to an LLM judge. Correlated errors undo that advantage. The LLM judges repeat nearly all of Jev's most confident errors, so a cascade replayed on the recorded verdicts lowers cost but gains at most 1.5 points over the best single judge with cross-fitted thresholds, and at most 2.0 even with oracle thresholds.
LLM-as-a-judge scales evaluation, but reasoning judges are slow and costly. We study JEV-as-a-Judge: evaluation with JEV, a decision-only judge that returns label probabilities instead of text, and whose confidence decides whether to accept its verdict or escalate to a reasoning judge. Against sixteen generative and reward-model judges, with blinded human adjudication, JEV comes within three points of GPT-6 wherever a verdict can be read off the text, at 0.36% of its fee and a 0.15-second median latency, and falls behind where the verdict must be derived, as in math, code, and logic. Its confidence marks this boundary. With a threshold frozen in advance, accepting confident verdicts and escalating the rest is 0.9 points more accurate than GPT-6 on 1,610 held-out pairs at 41% of its fee, and in a pre-specified live test on two new workloads the cascade matches GPT-6's accuracy exactly. Confidence routing weakens on style-adversarial pairs and reference-free prose; we close with a simple recipe for validating thresholds locally.