cs.LGSep 27, 2026

JET: Justification Evaluation in Transformer

Authors: Shenghao Ding

Organizations: Yet Another AI

Abstract

JET uses pretrained language and vision-language models to select among a finite set of answers without additional training. It evaluates candidate likelihoods directly and shares computation across candidates. Experiments on desktop CPUs and consumer GPUs assess decision accuracy and execution cost. Qwen3.6-35B-A3B achieves 87.48% accuracy on the full MMLU test set and 3.69 requests per second on a separately timed MMLU subset. The accuracy-throughput comparison covers model, hardware, and reasoning choices, with Jev as an external reference. Controlled execution experiments show 2.18-2.23-fold speedups from prefix reuse and cache management, and a 30.8% reduction in process time from input preparation optimizations, with unchanged outputs. Optional reasoning has a task-dependent accuracy-throughput trade-off. These results support local decision inference from existing models.

Figures & tables

Explore similar work

CardsList
  1. Evaluating and Benchmarking the System One Model Jev

    Sep 29, 2026Tobias Deußer, Lorenz Sparrenberg, Rafet SifaMultilingual Language Model EvaluationLLM Evaluation

  2. Visual Jev: Accurate and Efficient Decisions from Shared Visual Context

    Sep 22, 2026Guanxu Yu, Yuhang YaoVisual Question AnsweringEfficient Inference