JEV versus LLMs: Accuracy, Cost and Calibration on Seven Political Science Replications
Organizations: Leiden Institute for Area Studies, Leiden University. · Institute of Political Science, Leiden University.
Abstract
Large language models (LLMs) annotate and scale political text or constructs by generating text tokens. A new class of models, which TypeSafe markets as "System One" models, instead returns decisions and probability distributions across a user-supplied fixed answer set. A commercial model, JEV, is advertised as having a dramatic cost and speed advantage over traditional LLMs along with better calibrated decisions. As such, it might be useful for social scientists looking to quickly and cost-effectively annotate or scale large corpora of text and have a reliable indicator of a classifier's uncertainty. Yet, the accuracy of these claims and the broader model accuracy in social science text-based tasks are not yet established. In this paper, we do just that and hope to establish the suitability of JEV for social science tasks. We compare JEV with LLMs and human coders from published research, and with a current mid-tier commercial LLM (GPT-6 Luna) and an open-weight alternative (Qwen3.8-27B). We find that JEV matches, or comes close to, the capabilities of both LLMs in a variety of tasks. However, we find no cost advantage over GPT-6 Luna at OpenAI's batch prices. Further, we find that, when each question is asked once, JEV's probabilities are better calibrated than GPT-6 Luna's token probabilities, but not consistently better than Qwen3.8-27B's. We conclude that unless researchers have a need for speed, JEV's only obvious advantage is ease of parsing the underlying choice probabilities.
Figures & tables
| JEV compared with | ||||||
|---|---|---|---|---|---|---|
| Study | Data | Test replicated | Human reference | published models | GPT-6 Luna | Qwen 27B |
| Annotation | ||||||
| Gilardi et al. (2023) | Tweets on content moderation, 2021 and 2023 | Relevance (relevant or not), four runs; accuracy | Trained annotators; Mechanical Turk workers as a comparator | Better than ChatGPT | Worse | Slightly better (2021); worse (2023) a |
| Ornstein et al. (2025) | 945 tweets on two Supreme Court decisions | Sentiment (positive, neutral, negative); correlation of the probability-based score | Mean of three expert ratings | Worse than GPT-3 and GPT-4; better than Twitter RoBERTa and Naive Bayes | Worse | Worse |
| Brandt et al. (2026) | 37,709 Global Terrorism Database incidents; 322 BBC news articles | Attack type, nine categories, names only; accuracy and macro F1. BBC news: conflict F1 | GTD coders’ attack type | Better than prompted LLMs; worse than fine-tuned ConfliBERT and ConflLlama. BBC: matches Llama 3.1, worse than ConfliBERT | Better acc.; slightly better F1. BBC: better | Better acc. and F1. BBC: better |
| Weidmann et al. (2026) | 53 V-Dem indicators, 171 countries, 2023 (recall) | Ordinal indicator coded from the country name; correlation and exact agreement | V-Dem expert codes | Better than GPT-4o and Llama-3.1 70B | Better | Better |
| $ per 1,000 decisions | Seconds per decision | Decisions per second | ||||||||
| Application | JEV | Luna (Batch) | Luna (Std.) | Qwen | JEV | Luna | Qwen | JEV | Luna | Qwen |
| Gilardi et al. (2023) | 0.017 | 0.017 | 0.034 | 0.085 | 0.09 | 0.77 | 0.44 | 156 | 18 | 31 |
| Ornstein et al. (2025) | 0.024 | 0.012 | 0.025 | 0.128 | 0.26 | 0.75 | 0.39 | 47 | 18 | 30 |
| Brandt et al. (2026) | 0.027 | 0.018 | 0.035 | 0.139 | 0.27 | 1.04 | 0.87 | 51 | 14 | 14 |
| Weidmann et al. (2026) a | 0.020 | 0.012 | 0.024 | 0.058 | 0.26 | 0.75 | 0.41 | 45 | 17 | 31 |
| Di Leo et al. (2025) | 0.018 | 0.008 | 0.016 | 0.042 | 0.25 | 0.78 | 0.43 | 46 | 20 | 28 |
Appendix figures & tables12 assets
Supplementary material from the paper’s appendix.
Appendix
| Probability of choice | Pairs | Flip rate | Position bias |
|---|---|---|---|
| <.50 | 5,537 | 0.3655 | 0.1104 |
| .50–.70 | 2,984 | 0.1146 | 0.0721 |
| .70–.85 | 2,881 | 0.0139 | 0.0263 |
| .85–.95 | 3,862 | 0.0003 | 0.0079 |
| >.95 | 12,103 | 0.0000 | 0.0005 |
| Annotator | Accuracy | Intercoder agreement |
|---|---|---|
| GPT-6 Luna, single run | 91.1 | — |
| JEV 1.13, single run | 89.0 | — |
| Qwen3.8-27B, single run | 88.2 | — |
| JEV 1.13, four runs, unanimity required | 88.2 | 98.7 |
| ChatGPT (temp 1) | 72.8 | 92.0 |
| Mechanical Turk | 71.2 | 79.2 |
| Measure | Kind | Scoring rule | with experts |
| Qwen3.8-27B (few-shot) | generative | positive negative | 0.841 |
| GPT-6 Luna (few-shot) | generative | positive negative | 0.816 |
| GPT-3 (few-shot) | generative | principal component | 0.807 |
| GPT-6 Luna (zero-shot) | generative | positive negative | 0.791 |
| GPT-4 (few-shot) | generative | positive negative | 0.789 |
| JEV (few-shot) | constrained output | positive negative | 0.740 |
| Scorer | Mode | Expert pairs | Other models’ majority |
|---|---|---|---|
| Llama 3.1 405B | pairwise | 0.815 | 0.949 |
| Gemma 3 27B | pairwise | 0.813 | 0.928 |
| JEV | pairwise | 0.796 | 0.910 |
| GPT-4o | pairwise | 0.817 | 0.906 |
| GPT-4o mini | pairwise | 0.824 | 0.899 |
| Gemma 3 4B | pairwise | 0.739 | 0.875 |
| Attack type | Support | Confli- BERT | ConflLlama 8-bit | JEV names | JEV codebook | Qwen 3.5 names | Qwen 3.5 codebook | GPT-6 Luna names | GPT-6 Luna codebook | Qwen 27B names | Qwen 27B codebook |
|---|---|---|---|---|---|---|---|---|---|---|---|
| Assassination | 2,990 | 0.74 | 0.59 | 0.29 | 0.54 | 0.02 | 0.38 | 0.48 | 0.66 | 0.28 | 0.75 |
| Armed Assault | 9,079 | 0.81 | 0.75 | 0.90 | 0.94 | 0.94 | 0.89 | 0.79 | 0.89 | 0.88 | 0.91 |
| Bombing/Explosion | 14,508 | 0.97 | 0.92 | 0.91 | 0.81 | 0.84 | 0.65 | 0.92 | 0.86 | 0.85 | 0.73 |
| Hijacking | 154 | 0.64 | 0.54 | 0.60 | 0.70 | 0.44 | 0.58 | 0.38 | 0.50 | 0.55 | 0.56 |
| Hostage taking, barricade | 230 | 0.38 | 0.27 | 0.33 | 0.46 | 0.26 | 0.37 | 0.23 | 0.46 | 0.29 | 0.41 |
| Hostage taking, kidnapping | 3,495 | 0.88 | 0.86 | 0.75 | 0.82 | 0.77 | 0.81 | 0.71 | 0.81 | 0.69 | 0.71 |
| Measure | N | Precision | Recall | F1 | 95% CI | Accuracy |
|---|---|---|---|---|---|---|
| GPT-4o mini (paper) | 460 | 0.835 | 0.845 | 0.840 | [0.797, 0.876] | 0.824 |
| Gemma 3 27B (paper) | 460 | 0.821 | 0.841 | 0.831 | [0.787, 0.869] | 0.813 |
| GPT-4o (paper) | 460 | 0.852 | 0.805 | 0.828 | [0.780, 0.869] | 0.817 |
| JEV (Noul) | 460 | 0.812 | 0.829 | 0.821 | [0.772, 0.858] | 0.802 |
| Llama 3.1 405B (paper) | 459 | 0.877 | 0.769 | 0.820 | [0.770, 0.859] | 0.815 |
| JEV (rev) | 460 | 0.799 | 0.841 | 0.819 | [0.773, 0.858] | 0.798 |
| Test | Statistic | GPT-6 Luna | JEV | Difference | Reading |
|---|---|---|---|---|---|
| Tweet relevance, 2021 | accuracy (%) | 91.1 | 89.0 | 2.1 [0.7, 3.5] | above |
| Tweet relevance, 2023 | accuracy (%) | 84.2 | 76.2 | 8.0 [4.1, 11.9] | above |
| Tweet sentiment, few-shot | with experts | 0.816 | 0.740 | 0.076 [0.039, 0.112] | above |
| GTD attack type, names | accuracy | 0.709 | 0.715 | 0.006 [ 0.009, 0.003] | below |
| GTD attack type, names | macro F1 | 0.544 | 0.555 | 0.011 [ 0.025, 0.002] | level |
| BBC conflict | conflict F1 | 0.139 | 0.321 | 0.182 [ 0.308, 0.050] | below |
| Test | Statistic | Qwen | JEV | Difference | Reading |
| Tweet relevance, 2021 | accuracy (%) | 88.2 | 89.0 | 0.8 [ 2.1, 0.4] | level |
| Tweet relevance, 2023 | accuracy (%) | 79.6 | 76.2 | 3.4 [0.2, 6.8] | above a |
| Problem frame, 2021 | accuracy (%) | 92.4 | 92.4 | 0.0 [ 1.2, 1.2] | level |
| Problem frame, 2023 | accuracy (%) | 69.3 | 61.4 | 7.8 [1.2, 14.5] | above |
| Solution frame, 2021 | accuracy (%) | 81.4 | 84.7 | 3.4 [ 4.4, 2.3] | below |
| Solution frame, 2023 | accuracy (%) | 96.5 | 98.2 | 1.8 [ 4.1, 0.0] | level |
| Macro F1 | Unknown | |||
| Model | Accuracy | Nine types | Known types | recall |
| Fine-tuned on GTD labels (published) | ||||
| ConfliBERT | 0.840 | 0.744 | 0.775 | 0.550 |
| ConflLlama, 8-bit | 0.765 | 0.647 | 0.687 | 0.387 |
| ConflLlama, 4-bit | 0.729 | 0.571 | 0.608 | 0.323 |
| Prompted with the codebook | ||||
| Coder | Correlation | Difference | Exact | Changed | Within one | Missing |
|---|---|---|---|---|---|---|
| JEV, most probable level | 0.68 [0.66, 0.70] | 0.16 | 54 | 36 | 92 | 0 |
| JEV, expected level | 0.75 [0.73, 0.76] | 0.20 | 0 | |||
| GPT-6 Luna | 0.63 [0.61, 0.64] | 0.09 | 52 | 38 | 91 | 0 |
| Qwen3.8-27B | 0.59 [0.57, 0.61] | 0.17 | 51 | 35 | 89 | 0 |
| Qwen3.8-27B, expected level | 0.67 [0.66, 0.69] | 0.18 | 0 | |||
| GPT-4o | 0.63 [0.62, 0.65] | 0.28 | 49 | 36 | 90 | 5 |
| CHES experts | TEV voters | CMP manifestos | |||||||
|---|---|---|---|---|---|---|---|---|---|
| Year | GPT-3.5 | JEV | GPT-6 Luna | GPT-3.5 | JEV | GPT-6 Luna | GPT-3.5 | JEV | GPT-6 Luna |
| 1979 | — | — | — | 0.75 | 0.92 [0.79, 0.97] | 0.93 | 0.49 | 0.57 [0.44, 0.69] | 0.51 |
| 1984 | 0.77 | 0.88 [0.81, 0.92] | 0.83 | 0.78 | 0.92 [0.84, 0.96] | 0.89 | 0.53 | 0.59 [0.46, 0.69] | 0.56 |
| 1989 | 0.77 | 0.88 [0.81, 0.92] | 0.84 | 0.74 | 0.82 [0.65, 0.91] | 0.79 | 0.60 | 0.59 [0.46, 0.69] | 0.59 |
| 1994 | 0.81 | 0.89 [0.83, 0.93] | 0.84 | 0.84 | 0.89 [0.79, 0.94] | 0.83 | 0.44 | 0.52 [0.43, 0.60] | 0.50 |
| 1999 | 0.75 | 0.88 [0.83, 0.91] | 0.84 | 0.77 | 0.88 [0.81, 0.93] | 0.85 | 0.44 | 0.49 [0.39, 0.57] | 0.48 |
| UK manifestos (experts) | EU speeches (crowd) | ||
| Economic | Social | Subsidies | |
| JEV | 0.97 | 0.92 | 0.81 |
| JEV, relevance-weighted | 0.97 | 0.92 | 0.86 |
| GPT-4o (published) | 0.98 | 0.94 | 0.93 |
| GPT-6 Luna | 0.98 | 0.95 | 0.95 |
| Qwen3.8-27B | 0.97 | 0.94 | 0.90 |