cs.CVOct 8, 2026

VEDJE: Video-Efficient Discriminative Joint Encoder for Scalable Video-Text Retrieval

Authors: Shahaf Wagner, Gabriele Serussi, Dan Ben Ami, Tomer Galanti, Chaim Baskin

Organizations: INSIGHT Lab, Ben-Gurion University of the Negev, Israel · Decart AI · Texas A&M University, College Station

Abstract

Finding the right video often requires distinguishing similar scenes in which different events occur. Joint matching improves retrieval, but processing rich video representations for each query is costly. VEDJE compresses features within sampled frames while keeping their representations separate in a reusable cache. Feature-change prediction supplies an auxiliary training signal that improves retrieval from the compressed cache without adding work at query time. On MSR-VTT, MSVD, DiDeMo, and ActivityNet, VEDJE improves R@1 over matched first-stage retrievers in both retrieval directions. On MSR-VTT, it reaches 59.8 text-to-video R@1 with a fine-tuned VideoCLIP-XL first stage. In the VideoPrism configuration, shrinking the per-video cache fourfold to 12 KiB preserves text-to-video recall within 0.2 points. These results show that accurate video search can operate on compact evidence, encoded once and reused as new queries arrive.

Figures & tables

Appendix figures & tables15 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. VideoSearch-R1: Iterative Video Retrieval and Reasoning via Soft Query Refinement

    Jul 1, 2026Seohyun Lee, Seoung Choi, Dohwan Ko +2Temporal Video GroundingTemporal IR

  2. R^3: Composed Video Retrieval via Reasoning-Guided Recalling and Re-ranking

    May 31, 2026Zixu Li, Yupeng Hu, Zhiheng Fu +3Cross-Modal RetrievalComposed Video Retrieval

  3. Concentrate After Imagination: Text-Conditioned Evidence Grounding for Partially Relevant Video Retrieval

    Sep 8, 2026Shuaiqi Cheng, Siyu You, Yanbi Wu +3Temporal Video GroundingInformation Retrieval