How much does the decision readout matter when video-derived textual evidence is held fixed? We evaluate Jev typed decisions and three Qwen readouts on a sparse development sample of 40 videos and 400 target anchors from UCF-Crime and XD-Violence, each presented as a summary and ordered captions. Each dataset contributes 20 source groups and 200 anchors, including only 10 and 37 positives, respectively. The original five-backend pilot requested 4,000 predictions; Jev Choice returned 776 valid responses out of 800 under the study's strict numerical policy, blocking its full-coverage quality comparison. On XD captions, Jev Noul achieved 75.99% average precision versus 48.47% for Qwen generated probability and 57.81% for the stronger local ordinal-likelihood expectation. The latter paired difference was 18.18 percentage points (95% source-group bootstrap interval 5.53-31.50). UCF did not show a corresponding advantage: caption ROC-AUC was 52.26% for Noul and 65.95% for ordinal likelihood. Both probability readouts had higher, hence worse, UCF Brier scores than the evaluation-prevalence reference of 0.0475. We additionally audit historical LAVAD scores at exactly matched anchors and distinguish response structure from numerical consistency. A binary-likelihood control is missing. These exploratory offline results characterize ranking, probability quality and interface failures; they establish neither a causal typed-interface benefit nor general superiority, calibration or end-to-end acceleration.
Figures & tables
View
Readout
Coverage
Primary
95% interval
Sum
Noul
200/200
47.95
[25.10, 74.07]
Sum
Qwen generated prob.
200/200
56.21
[41.86, 73.21]
Sum
Qwen ordinal
200/200
56.42
[31.91, 78.69]
Sum
Qwen short
200/200
54.92
[39.47, 74.83]
Sum
Choice
198/200
—
—
Cap
Noul
200/200
52.26
[24.39, 83.34]
Table 1: UCF sampled-anchor ROC-AUC (%) and 95% group intervals. Coverage is valid/requested. All rows cover the same target inventory; Cache is one historical reference, not a view-specific rerun. 1,997 valid ranking replicates / 2,000.
View
Readout
Coverage
Primary
95% interval
Sum
Noul
200/200
61.49
[36.02, 84.05]
Sum
Qwen generated prob.
200/200
48.73
[26.33, 70.55]
Sum
Qwen ordinal
200/200
56.12
[33.20, 78.13]
Sum
Qwen short
200/200
53.80
[31.66, 71.65]
Sum
Choice
191/200
—
—
Cap
Noul
200/200
75.99
[47.46, 92.48]
Table 2: XD sampled-anchor standard AP (%) and 95% group intervals. Coverage is valid/requested. All rows cover the same target inventory; Cache is one historical reference, not a view-specific rerun. 2,000 valid ranking replicates / 2,000.
Data
View
Readout / reference
Brier
NLL
ECE
UCF
Sum
Noul
0.0904
0.3032
0.0898
UCF
Sum
Qwen generated prob.
0.0751
0.4239
0.0632
UCF
Cap
Noul
0.0933
0.3113
0.0903
UCF
Cap
Qwen generated prob.
0.1154
1.1993
0.1109
UCF
Both
p=0
0.0500
1.7269
—
UCF
Both
p=π (descriptive)
0.0475
0.1985
—
Table 3: Event-probability quality; every cell covers 200/200 anchors. Lower Brier/NLL is better. Constants apply to both views, not extra samples. The evaluation-prevalence reference is post hoc and label-informed. ECE for constants is intentionally omitted.
Data
View
Coverage
Reference
Ref. score
Choice score
Choice − Ref.
95% interval
UCF
Sum
198/200
O
56.60
54.34
−2.26
[ −14.32 , 10.79]
UCF
Sum
198/200
N
48.14
54.34
6.20
[ −4.68 , 18.77]
UCF
Cap
194/200
O
66.25
62.61
−3.64
[ −16.69 , 8.27]
UCF
Cap
194/200
N
52.91
62.61
9.70
[ −2.25 , 26.29]
XD
Sum
191/200
O
56.54
60.26
3.72
[ −10.63 , 18.44]
XD
Sum
191/200
N
60.74
60.26
−0.48
[ −5.66 , 7.69]
Table 4: Matched-subset diagnostic only: Choice expectation and each reference are recomputed on the identical valid subset. O = Qwen ordinal expectation; N = Jev Noul. Scores are UCF AUC or XD AP (%); differences are percentage points. Intervals use 1,997 UCF / 2,000 XD valid replicates.
Readout
Valid
Median
p95
Wall sum
Choice
776/800
0.3181
0.3720
264.70
Noul
800/800
0.3139
0.3664
261.72
Qwen generated prob.
800/800
0.5775
0.7681
591.41
Qwen ordinal
800/800
1.3215
1.6133
1201.36
Qwen short
800/800
0.6379
0.7961
636.87
Table 5: Pilot deployment observations in seconds. Medians/p95 include valid requests only; wall sums include all outcomes and backend logging, but are not elapsed run time.
Appendix figures & tables7 assets
Supplementary material from the paper’s appendix.
Appendix
Data
View
Readout
n
Min
Median
p95
Max
UCF
Sum
C
200
709
732.0
762.00
774
UCF
Sum
N
200
371
394.0
424.00
436
UCF
Sum
G
200
139
162.0
192.00
204
UCF
Sum
O/S
200
410
433.0
463.00
475
UCF
Cap
C
200
939
987.5
1020.05
1030
UCF
Cap
N
200
601
649.5
682.05
692
Appendix
Table 6: Actual input token counts (200 requests per cell). O and S have identical prompts/counts. C/N are provider-reported and cannot be interpreted as Qwen token counts.
Data
View
Readout
Coverage
AUC
AP
PR trap.
UCF
Sum
C
198/200
—
—
—
UCF
Sum
N
200/200
47.95
7.02
5.40
UCF
Sum
G
200/200
56.21
5.74
4.95
UCF
Sum
O
200/200
56.42
6.62
5.65
UCF
Sum
O argmax
200/200
54.92
5.59
6.14
UCF
Sum
S
200/200
54.92
5.59
6.14
Appendix
Table 7: All original pilot metric cells; ROC-AUC, standard AP and trapezoidal PR area (%). Extra argmax rows preserve the original alternate aggregation. Missing full-coverage values are not matched-subset estimates.
Data
View
Contrast
Readout
n
Diff.
95% interval
Valid
UCF
Sum
C − S
scalar
198
−0.48
[ −10.06 , 8.89]
1997
UCF
Sum
C − S
arg
198
2.26
[0.26, 5.22]
1997
UCF
Sum
C − O
scalar
198
−2.26
[ −14.32 , 10.79]
1997
UCF
Sum
C − O
arg
198
2.26
[0.26, 5.22]
1997
UCF
Sum
O − S
scalar
200
1.50
[ −12.61 , 15.77]
1997
UCF
Sum
O − S
arg
200
0.00
[0.00, 0.00]
1997
Appendix
Table 8: All 44 original paired contrasts, retained without outcome selection. Differences use UCF AUC / XD AP percentage points. Each interval has 2,000 requested replicates; Valid counts non-single-class replicates. Rows involving C use its matched subset; others have full coverage, except C view pairs which require validity in both views. All comparisons have 20 groups.
Data
View
Contrast
n
Diff.
95% interval
Valid
UCF
Sum
N − G
200
−8.26
[ −24.28 , 5.86]
1997
UCF
Sum
N − O
200
−8.47
[ −28.04 , 10.11]
1997
UCF
Sum
N − L
200
−0.97
[ −15.00 , 11.85]
1997
UCF
Sum
O − L
200
7.50
[ −6.21 , 19.13]
1997
UCF
Sum
G − L
200
7.29
[ −9.78 , 24.28]
1997
UCF
Cap
N − G
200
−1.50
[ −33.97 , 34.17]
1997
Appendix
Table 9: Recomputed post-pilot ranking contrasts in percentage points. L is the same historical raw reference across both view comparisons, not two independent LAVAD runs. All have 20 groups and 2,000 requested replicates.
Data
View
Metric
Diff.
95% interval
Valid
UCF
Sum
BRIER
0.0153
[ −0.0012 , 0.0356]
2000
UCF
Sum
NLL
−0.1207
[ −0.4456 , 0.0651]
2000
UCF
Cap
BRIER
−0.0221
[ −0.0649 , 0.0061]
2000
UCF
Cap
NLL
−0.8880
[ −1.7734 , −0.2112 ]
2000
XD
Sum
BRIER
−0.0005
[ −0.0429 , 0.0385]
2000
XD
Sum
NLL
−0.0138
[ −0.1184 , 0.0809]
2000
Appendix
Table 10: Exploratory Noul minus Qwen generated-probability differences on the original probability scale, not percentage points. Both methods cover 200/200 anchors; all 2,000 replicates are defined, in 20 groups.
First failure
Choice
Confidence
Sum
mass
1.0
0.8
0.99
choice not max
0.0
0.06
1.00
Appendix
Table 11: Actual numeric failure extracts. The full ordered distributions are shown immediately below.
Data
Prediction
AUC
AP
PR trap.
UCF
Raw
72.7904
18.2100
19.5168
UCF
Original refinement
80.2757
27.1911
27.0773
XD
Raw
80.6145
53.0951
57.2054
XD
Original refinement
85.3642
62.0045
62.0058
Appendix
Table 12: Historical full-frame replay, distinct from every pilot table. Values (%) are retained at audit precision; these are cached author predictions, not new model runs.
Video text-based visual question answering (Video TextVQA) aims to answer questions by reasoning over visual textual content appearing in videos. Despite the strong multimodal video understanding capabilities of recent Video-LLMs, their performance on existing Video TextVQA benchmarks remains limited. To better understand this gap, we conduct an upper-bound analysis through frame-wise question answering, counting a sample as correct if any frame yields the right answer, which significantly outperforms direct video-based inference and reveals a substantial performance gap. The results suggest that the primary bottleneck lies in the localization of key question-relevant evidence, rather than in reasoning capacity itself. Building on this insight, we propose a question-guided agent framework that explicitly anchors the relevant keyframes before answering. The approach operates effectively in a training-free setting and consistently surpasses direct video inference. With additional supervised fine-tuning (SFT) and reinforcement learning (RL), it achieves an average improvement of +12.12 in accuracy and +11.15 in ANLS across benchmarks, establishing new state-of-the-art results. Our study underscores the critical role of explicit keyframe anchoring for advancing Video TextVQA. The code will be publicly released.
Haibin He, Maoyuan Ye, Jing Zhang +2
School of Computer Science, National Engineering Research Center for Multimedia Software, Institute of Artificial Intelligence, and Hubei Key Laboratory of Multimedia and Network Communication Engineering, Wuhan University, Wuhan, Hubei, China
Many decision-making tasks, where both accuracy and efficiency matter, still require human supervision. For example, tasks like traffic officers reviewing hour-long dashcam footage or researchers screening conference videos can benefit from concise summaries that reduce cognitive load and save time. Yet current vision-language models (VLMs) often produce verbose, redundant outputs that hinder task performance. Existing video caption evaluation depends on costly human annotations and overlooks the summaries' utility in downstream tasks. We address these gaps with Video-to-text Information Bottleneck Evaluation (VIBE), an annotation-free method that scores VLM outputs using two metrics: grounding (how well the summary aligns with visual content) and utility (how informative it is for the task). VIBE selects from randomly sampled VLM outputs by ranking them according to the two scores to support effective human decision-making. Human studies on LearningPaper24, SUTD-TrafficQA, and LongVideoBench show that summaries selected by VIBE consistently improve performance-boosting task accuracy by up to 61.23% and reducing response time by 75.77% compared to naive VLM summaries or raw video.
LLM-as-a-judge scales evaluation, but reasoning judges are slow and costly. We study JEV-as-a-Judge: evaluation with JEV, a decision-only judge that returns label probabilities instead of text, and whose confidence decides whether to accept its verdict or escalate to a reasoning judge. Against sixteen generative and reward-model judges, with blinded human adjudication, JEV comes within three points of GPT-6 wherever a verdict can be read off the text, at 0.36% of its fee and a 0.15-second median latency, and falls behind where the verdict must be derived, as in math, code, and logic. Its confidence marks this boundary. With a threshold frozen in advance, accepting confident verdicts and escalating the rest is 0.9 points more accurate than GPT-6 on 1,610 held-out pairs at 41% of its fee, and in a pre-specified live test on two new workloads the cascade matches GPT-6's accuracy exactly. Confidence routing weakens on style-adversarial pairs and reference-free prose; we close with a simple recipe for validating thresholds locally.