cs.CLOct 5, 2026

JEV versus LLMs: Accuracy, Cost and Calibration on Seven Political Science Replications

Authors: Steven Denney, Matthew DiGiuseppe

Organizations: Leiden Institute for Area Studies, Leiden University. · Institute of Political Science, Leiden University.

Abstract

Large language models (LLMs) annotate and scale political text or constructs by generating text tokens. A new class of models, which TypeSafe markets as "System One" models, instead returns decisions and probability distributions across a user-supplied fixed answer set. A commercial model, JEV, is advertised as having a dramatic cost and speed advantage over traditional LLMs along with better calibrated decisions. As such, it might be useful for social scientists looking to quickly and cost-effectively annotate or scale large corpora of text and have a reliable indicator of a classifier's uncertainty. Yet, the accuracy of these claims and the broader model accuracy in social science text-based tasks are not yet established. In this paper, we do just that and hope to establish the suitability of JEV for social science tasks. We compare JEV with LLMs and human coders from published research, and with a current mid-tier commercial LLM (GPT-6 Luna) and an open-weight alternative (Qwen3.8-27B). We find that JEV matches, or comes close to, the capabilities of both LLMs in a variety of tasks. However, we find no cost advantage over GPT-6 Luna at OpenAI's batch prices. Further, we find that, when each question is asked once, JEV's probabilities are better calibrated than GPT-6 Luna's token probabilities, but not consistently better than Qwen3.8-27B's. We conclude that unless researchers have a need for speed, JEV's only obvious advantage is ease of parsing the underlying choice probabilities.

Figures & tables

Appendix figures & tables12 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. JEV vs. LLMs as Rubric Judges: Cheaper, Faster, and Wrong in the Same Places

    Sep 24, 2026Delip Rao, Chris Callison-BurchLarge Language Model Judges

  2. JEV-as-a-Judge: Accept When Confident, Escalate When Unsure

    Sep 22, 2026Yubo Li, Yidi Miao, Ramayya Krishnan +1Large Language Model JudgesJudges