Video Evidence Indexing: Learning Where to Look from Video Previews for Token-Budgeted Long-Video Question Answering
Abstract
Long-video question answering is limited by the high cost of visual tokens and by the fixed context width of current VLMs. A long-video question may require broad temporal coverage, but the answer is often supported by only a compact set of moments. To locate these moments efficiently, we propose token-budgeted Video Evidence Indexing (VEI): given a dense low-resolution Video Preview, the model constructs a compact high-resolution Evidence Set for final reasoning. We treat VEI as a policy that must jointly solve \textit{evidence localization}, which finds question-relevant moments, and \textit{budget planning}, which decides where to spend the limited high-resolution frame budget. We implement this idea with an inference pipeline: the Video Preview provides cheap global coverage, Video Evidence Indexing constructs the Evidence Set, and Answer Generation combines both inputs for final VQA. To address missing frame-level supervision, we adopt privileged self-distillation, where an answer-aware teacher guides the normal test-time policy on student-generated indexing traces. We explore previews at 1, 6, 12, and 24 visual tokens per frame, training a single policy that supports all four resolutions. Experiments show that Video Evidence Indexing improves accuracy under limited visual budgets, and self-distillation further improves both QA accuracy and temporal evidence localization.
Figures & tables
| Method | Context Limit | Preview Size | VideoMME w/o sub. | LVBench | MLVUMCQ | Average | ||||
| Short | Medium | Long | Overall | Train | Test | Score | % | |||
| Use 100% Context Width Across All Benchmarks | ||||||||||
| Qwen3-VL-8B | ||||||||||
| Use 5.9% Context Width Across All Benchmarks | ||||||||||
| Uniform FS | ||||||||||
| Training-Free VEI | 72.1\;{\color[rgb]{0,0,1}(-7.5)} | |||||||||
| Model | Size | Token Budget | VideoMME | LVBench | MLVU | AVG | |
| Long | Overall | ||||||
| (a) Proprietary Models | |||||||
| GPT-4o | – | 65.3 | 71.9 | 64.4 | 64.6 | 67.0 | |
| Gemini-2.5-Pro | – | – | 87.0 | 69.2 | 81.2 | 79.1 | |
| (b) Open-Source / Public Pretrained Models | |||||||
| LLaVA-OneVision ( Li et al., 2024 ) | 7B | 46.7 | 58.2 | 26.9 | 50.5 | 45.2 | |
Appendix figures & tables7 assets
Supplementary material from the paper’s appendix.
Appendix
| Selector | Comparable questions | Base FS | + EL | + EL + BP (VEI) |
| Training-Free VEI | 220 / 502 | 70.0 | 66.4 | 74.6 |
| Self-Distilled VEI | 218 / 502 | 78.9 | 78.0 | 80.7 |
| Method | Video Processed | Average Accuracy All Benchmarks | Estimated Cost All Benchmarks |
| Qwen3-VL-8B | 65.4 | \sim\87.8$ | |
| Self-Distilled VEI | 66.9 | \sim\19.1$ |
| Question | Outcome | Diagnosis |
| 61 | Hit, wrong | The selected evidence reaches the annotated event but misses its later outcome. |
| 65 | Hit, wrong | The interaction is localized, but the answer misses the semantic detail that no bribe occurred. |
| 73 | Miss, wrong | Selection concentrates on unrelated moments and misses the annotated event. |
| 78 | Miss, correct | The answer is correct despite missing the annotated window; the preview or alternative evidence may provide support. |
| Method | Context Limit | VideoMME w/o sub. | LVBench | MLVUMCQ | Average Score ( ) | |||
| Short | Medium | Long | Overall | Train | Test | |||
| Use 5.9% Context Width Across All Benchmarks | ||||||||
| Training-Free VEI | ||||||||
| w/o Evidence Set | ||||||||
| Use 7.8% Context Width Across All Benchmarks | ||||||||
| Training-Free VEI | ||||||||
| Parameter | SD |
| Learning Rate | |
| Effective Batch Size | |
| LoRA Rank ( ) | |
| LoRA Alpha ( ) | |
| LoRA Target Modules | q_proj, k_proj, v_proj, o_proj, gate_proj, up_proj, down_proj |
| Max Completion Length |