Have I Seen Enough? Frozen Video-Language Models Encode Evidence Readiness
Organizations: INSIGHT Lab, Ben-Gurion University of the Negev, Israel · Ben-Gurion University of the Negev, Israel
Abstract
Streaming video-language models must decide not only what to answer, but whether the evidence needed for the current question has arrived. Existing systems learn that decision as a separate trigger; we ask whether an unmodified model already computes it. We show that frozen VideoLLMs carry a linearly readable evidence-readiness signal, labelled from timestamped evidence rather than from model output. It decodes in all seven models of a shared byte-identical evaluation (AUROC 0.733-0.905 under the strictest not-ready sampling, where a fitted clock is near chance), and a probe fitted without any of a benchmark family's footage still reads that family. It is question-conditioned: on byte-identical windows, changing only the question reverses the readout on 66.1% of pairs, while every question-blind control is at chance by construction. The model can answer incorrectly and still encode readiness: AUROC remains 0.722 among wrong answers. Readiness also beats uncertainty estimators and their supervised combination on latency-matched answer selection, and tracks independent human judgments more closely than confidence. Released streaming triggers are also linear readouts, yet a trained trigger read on its own base model's activations is approximately orthogonal to readiness and decodes it far less accurately than a probe. We turn the readout into Readiness Gating, an answer-timing policy that improves accuracy by up to +9.75 pp at matched video duration with negligible computational overhead. How much it gains varies with the accuracy headroom the task makes available: across 26 configurations the gain tracks that headroom, and an intervention that moves it over identical pixels moves the gain with it.
Figures & tables
| source | geometry | label | units | shuffle control | confidence | visual layer- | readiness AUROC |
|---|---|---|---|---|---|---|---|
| HD-EPIC † | composite | cue present | 520 prefixes | ||||
| HD-EPIC † | matched pair | availability | 4,800 windows | ||||
| NExT-GQA | quad | containment | 1,208 quads | ||||
| NExT-GQA | prefix | onset | 5,580 windows | ||||
| ST-Evidence | prefix | onset | 6,546 windows | ||||
| ST-Evidence-Gen | prefix | onset | 4,346 prefixes |
| measurement | Oops! | ActivityNet-Captions |
|---|---|---|
| readiness AUROC on evaluated windows | ||
| elapsed-time AUROC on the same windows | ||
| selection accuracy, answering at end of video | ||
| selection accuracy, confidence | ||
| selection accuracy, Readiness Gating (ours) | ||
| gating advantage over supervised combination | pp | pp |
| HERBench shot-ID | ActivityNet, same-action | |||
| monitor | readiness AUROC | selection acc. | readiness AUROC | selection acc. |
| answer at end (no selection) | — | — | ||
| evidence oracle | — | — | ||
| Readiness Gating (ours) | ||||
| confidence | ||||
| elapsed time | — | — | ||
Appendix figures & tables26 assets
Supplementary material from the paper’s appendix.
Appendix
| geometry | shape | frames | label | what it is for |
|---|---|---|---|---|
| growing prefix | , evidence-relative ends | onset, | graded accumulation; earliest evidence | |
| recent window | , fixed span, overlap | containment | first evidence-agnostic ladder | |
| containment window | s span, free s stride | @ fps | containment | the reference format; model axis, utility |
| short window | s span, s stride | @ fps | containment | point-onset footage |
| matched pair | rows per item, all else fixed | @ fps | availability | availability vs. correctness; causal work |
| crossed quad | 2 questions 2 windows | @ fps | containment | question-conditioning; visual controls at |
| label | definition | used for |
|---|---|---|
| containment | coverage contains the evidence hull; null on partial | probe fitting |
| deployed | the same, partial folded into not-ready | policy evaluation |
| onset | any delivered frame inside an evidence segment | threshold sensitivity |
| scored against the benchmark’s evidence label | NExT-GQA | ActivityNet-Cap. | Oops! |
|---|---|---|---|
| AUROC, readiness probe | |||
| AUROC, independent rater verdict | |||
| rater accuracy | |||
| rater | |||
| rater sensitivity / specificity | |||
| screens ( ) |
| scored against the raters’ readiness verdict | NExT-GQA | ActivityNet-Cap. | Oops! |
| AUROC, readiness probe | |||
| AUROC, the model’s own confidence | |||
| AUROC, an external LLM asked directly | |||
| model accuracy, rater-ready screens | |||
| model accuracy, rater-not-ready screens | |||
| screens ( ), probe and confidence |
| rater answer accuracy | pre (far) | pre (near) | partial | first ready | post | overall |
|---|---|---|---|---|---|---|
| NExT-GQA | ||||||
| ActivityNet-Cap. | ||||||
| Oops! |
| evaluation | held-out unit | fitting and scoring |
|---|---|---|
| controlled-composite headline | cue object | Nested within-pool fitting; evaluation on held-out objects. |
| controlled-composite visual null | pair-preserving group | Both questions on identical frames share a fold; distinct from the headline protocol. |
| natural within-benchmark probes, including the model axis | video | Video-grouped inner selection and outer evaluation; matched pairs and crossed quads stay together. |
| held-out-family transfer | footage family | Selection on source families only; the target family’s footage is excluded from fitting. |
| external transfer | disjoint source and target videos | Source scaler, layer and probe are frozen on the target; a target-fitted probe is a separate comparison. |
| layer | |||||||||
|---|---|---|---|---|---|---|---|---|---|
| evaluation | probe | final | vision | option | shuffle | clock | conf. | resid. drop | over conf. |
| NExT-GQA | |||||||||
| ST-Evidence-MCQ | |||||||||
| ST-Evidence-Gen | |||||||||
| the streaming cohort | |||||||||
| HERBench shot-ID | |||||||||
| pool | footage | evidence timestamps | items | rows | source groups | readiness | model acc. |
|---|---|---|---|---|---|---|---|
| controlled composites, crossed | HD-EPIC egocentric, P01–P09 | cue presence by construction, 17-object library | 260 (520 prefixes) | 1,040 | 31 videos, 90 clips | 0.5 by constr. | – |
| matched present / absent windows | HD-EPIC egocentric, P01–P09 | official fine-grained action segments | 2,400 pairs | 4,800 | 142 videos | 0.5 by constr. | – |
| fine-grained action VQA | HD-EPIC egocentric | the benchmark’s own items | 1,200 | – | 86 videos | – | – |
| NExT-GQA prefixes | NExT-QA | human temporal grounding | 1,000 | 5,695 | 473 videos | 3,600 / 2,095 | 0.729 five-opt., 0.695 six-opt. |
| NExT-GQA crossed quads | NExT-QA | human temporal grounding | 1,208 quads (1,196 intact) | 4,808 | 277 videos | 0.5 by constr. | – |
| ST-Evidence-MCQ prefixes | NExT-QA | independent human evidence | 1,298 | 6,756 | 348 videos | 4,400 / 2,356 | 0.704 |
| evaluation | questions | source of the items |
|---|---|---|
| HD-EPIC controlled composites | constructed | synthetic cues composited onto real footage |
| HD-EPIC matched present / absent | constructed | from the official fine-grained action segments |
| HD-EPIC fine-grained action VQA | shipped | the benchmark’s own items, read verbatim |
| NExT-GQA | shipped | NExT-QA’s questions with human grounding |
| ST-Evidence-MCQ | shipped | the benchmark’s own QA and evidence |
| ST-Evidence-Gen | shipped | the benchmark’s own items, two item filters applied |
| model | HuggingFace id | vision tower | language backbone | dec. lay. | how frames are fed | probed layers |
|---|---|---|---|---|---|---|
| Qwen3-VL-8B-Instruct | Qwen/Qwen3-VL-8B-Instruct | Qwen-ViT | Qwen3-8B | 36 | native video path, video metadata | 0, 14, 18, 23, 29, 32 |
| Qwen3-VL-2B-Instruct | Qwen/Qwen3-VL-2B-Instruct | Qwen-ViT | Qwen3-2B | – | native video path | 0, 11, 14, 18, 22, 25 |
| Qwen3-VL-32B-Instruct | Qwen/Qwen3-VL-32B-Instruct | Qwen-ViT | Qwen3-32B | – | native video path | 0, 26, 32, 42, 51, 58 |
| InternVL3.5-8B [ Wang et al., 2025b ] | OpenGVLab/InternVL3_5-8B-HF | InternViT | Qwen3-8B | 36 | multi-image, one px tile per frame | 0, 14, 18, 23, 29, 32 |
| LLaVA-OneVision-7B [ Li et al., 2025 ] | llava-hf/llava-onevision-qwen2-7b-ov-hf | SigLIP [ Zhai et al., 2023 ] | Qwen2-7B | 28 | native video path, no metadata | by depth fraction |
| LLaVA-OneVision-1.5-8B [ An et al., 2025 ] | lmms-lab/LLaVA-OneVision-1.5-8B-Instruct | – | Qwen3-8B | – | – | 0, 14, 18, 23, 29, 32 |
| geometry | frames | rate | span | stride | decode resolution |
|---|---|---|---|---|---|
| controlled composites | 16 | 2 fps | fixed 8 s | – | 448 px long side |
| growing prefixes | 32 | – | evidence-relative ends | 448 px long side | |
| containment windows | 32 | 2 fps | 16 s | free 2 s | – |
| short windows | 8 | 4 fps | 2 s | 0.5 s | – |
| sparse bursts | 16 | 2–3 fps | 3–6 bursts of 1–2 s | one block | – |
| pool | dropped | of | rate |
|---|---|---|---|
| NExT-GQA prefixes | 115 prefixes (21 items entirely) | 5,695 | 2.0% |
| ST-Evidence-MCQ prefixes | 210 prefixes | 6,756 | 3.1% |
| ST-Evidence-Gen prefixes | 5 prefixes | 4,346 | 0.1% |
| matched present / absent windows | 0 | 4,800 | 0% |
| model-axis short windows | 0 | – | 0% |
| Oops! short windows | – | – | 0.3% |
| pool | geometry | label | rows / clusters | probe AUROC | shuffle | L @vision | clock |
|---|---|---|---|---|---|---|---|
| controlled composites | controlled composite | cue presence | 1,040 / 17 obj. | – | |||
| matched windows | matched pair | availability | 4,800 / 142 | ||||
| the same, InternVL3.5-8B | matched pair | availability | 4,800 / 142 | ||||
| NExT-GQA prefixes | growing prefix | onset | 5,580 / 458 | ||||
| ST-Evidence-MCQ prefixes | growing prefix | onset | 6,546 / 338 | ||||
| ST-Evidence-Gen prefixes | growing prefix | 4,346 / 1,529 |
| paired drop on byte-identical frames, pairs | |||||
|---|---|---|---|---|---|
| model | language backbone | identical question | paraphrase | distinct question | register-matched |
| Qwen3-VL-2B | Qwen3 | ||||
| Qwen3-VL-8B | Qwen3 | ||||
| Qwen3-VL-32B | Qwen3 | ||||
| LLaVA-OV-1.5-8B | Qwen3 | ||||
| GLM-4.1V-9B | GLM-4 | ||||
| pool | readiness AUROC | residualized on conf. | drop | confidence alone | increment |
|---|---|---|---|---|---|
| controlled composites, factorial | |||||
| controlled composites, crossed | |||||
| matched windows | |||||
| the same, InternVL3.5-8B | |||||
| NExT-GQA prefixes | |||||
| ST-Evidence-MCQ prefixes |
| Oops! windows | ActivityNet | ||||
| signal | what it is | acc. | probe this | acc. | probe this |
| sc_agree | share of sampled answers, each with a rationale, agreeing with the mode | ||||
| sc_neg_entropy | negative entropy of that sampled answer distribution | ||||
| se_neg | semantic entropy: open-ended answers clustered by bidirectional entailment | ||||
| se_neg_lennorm | the same, length-normalized | ||||
| se_neg_loose | the same, one-direction clustering | ||||
| rows | ready | not ready | AUROC | |
|---|---|---|---|---|
| 2,309 | 1,877 | 432 | ||
| 1,121 | 667 | 454 | ||
| 379 | 174 | 205 | ||
| 211 | 81 | 130 | ||
| 156 | 51 | 105 | ||
| 75 | 24 | 51 |
| backbone | model | regime | AUROC | CI | L @final | L @vision | shuffle | clock track |
|---|---|---|---|---|---|---|---|---|
| Qwen3 | Qwen3-VL-2B | before | ||||||
| near | ||||||||
| after | ||||||||
| both | ||||||||
| Qwen3 | Qwen3-VL-8B | before | ||||||
| near |
| held-out family | transfer AUROC | CI | own AUROC | retention | L @vision | shuffle | clock |
| NExT-GQA | |||||||
| ST-Evidence-MCQ | |||||||
| ST-Evidence-Gen | |||||||
| ActivityNet | |||||||
| Oops! | |||||||
| mean retention, footage-group hold-out | |||||||
| rank | aligned (median) | random-rotation null | row-shuffled alignment | pairs beating own null |
|---|---|---|---|---|
| cell | base rate | carried-over | at | Platt at | oracle threshold | constant predictor |
|---|---|---|---|---|---|---|
| OVO-Bench (window) | ||||||
| Fine-grained Action | ||||||
| NExT-GQA ladder | ||||||
| ActivityNet ladder |
| system | head | shape | base model |
|---|---|---|---|
| MMDuet | relevance , informative | ea. | LLaVA-OneVision (Qwen2-7B) |
| Dispider | silent | Qwen2 | |
| VideoLLM-online | streaming-EOS logit | unembedding row | Llama-3-8B |
| quantity | Qwen3-VL-8B | InternVL3.5-8B |
|---|---|---|
| commitment excess over a magnitude-matched control | ||
| deferral-to-answer flips, treatment vs. control | vs. | vs. |
| exceeds all alternative directions | yes, | yes, same |
| dose–response Spearman over four doses | ||
| answer-content equivalence test | passes | passes |
| replicates on evidence-present-but-wrong | yes | yes |