VEDJE: Video-Efficient Discriminative Joint Encoder for Scalable Video-Text Retrieval
Organizations: INSIGHT Lab, Ben-Gurion University of the Negev, Israel · Decart AI · Texas A&M University, College Station
Abstract
Finding the right video often requires distinguishing similar scenes in which different events occur. Joint matching improves retrieval, but processing rich video representations for each query is costly. VEDJE compresses features within sampled frames while keeping their representations separate in a reusable cache. Feature-change prediction supplies an auxiliary training signal that improves retrieval from the compressed cache without adding work at query time. On MSR-VTT, MSVD, DiDeMo, and ActivityNet, VEDJE improves R@1 over matched first-stage retrievers in both retrieval directions. On MSR-VTT, it reaches 59.8 text-to-video R@1 with a fine-tuned VideoCLIP-XL first stage. In the VideoPrism configuration, shrinking the per-video cache fourfold to 12 KiB preserves text-to-video recall within 0.2 points. These results show that accurate video search can operate on compact evidence, encoded once and reused as new queries arrive.
Figures & tables
| (a) Published systems | ||
|---|---|---|
| Method | T2V R@1 | V2T R@1 |
| CLIP4Clip B/32 ( Luo et al., 2022 ) | 44.5 | 43.1 |
| X-CLIP B/16 ( Ma et al., 2022 ) | 49.3 | 48.9 |
| Video-ColBERT (SigLIP-B/16) ( Reddy et al., 2025 ) | 51.5 | – |
| CrossTVR-Large ( Dai et al., 2025 ) | 54.0 | 51.3 |
| UMT-L (ViT-L/16) ( Li et al., 2023b ) | 58.8 | 58.6 |
| Method | Backbone treatment | T2V R@1 | V2T R@1 |
|---|---|---|---|
| CLIP4Clip ( Luo et al., 2022 ) | fine-tuned | 46.4 | 45.4 |
| X-CLIP ( Ma et al., 2022 ) | fine-tuned | 49.3 | 48.9 |
| EERCF ( Tian et al., 2024a ) | fine-tuned | 49.9 | 47.8 |
| VEDJE-CLIP | frozen | 49.8 | 49.6 |
| Tokens | KiB/video | Direction | Without | With | Gain |
|---|---|---|---|---|---|
| 64 | 48 | T2V | +0.6 | ||
| 64 | 48 | V2T | +1.0 | ||
| 16 | 12 | T2V | +1.9 | ||
| 16 | 12 | V2T | +0.5 |
| Cache construction | T2V R@1 | V2T R@1 | |
|---|---|---|---|
| FrameSet-EDJE: mean pooling | No | ||
| FrameSet-EDJE: attention pooling | No | ||
| VEDJE: separate frame groups | No | ||
| VEDJE: separate frame groups | Yes |
| Scoring rule | T2V R@1 |
|---|---|
| Stage 1: VideoPrism | 50.1 |
| Trained MaxSim with residual prior | 51.9 |
| Joint encoder with residual prior (VEDJE) | 54.6 |
| Latency (ms) | T2V R@1 (%) | |||
|---|---|---|---|---|
| Candidates | 64 tokens | 16 tokens | 64 tokens | 16 tokens |
| 20 (default) | 5.79 | 5.82 | 54.6 | 54.4 |
| 50 | 6.21 | 6.12 | 53.9 | 54.7 |
| 100 | 7.83 | 6.33 | 54.4 | 54.6 |
| 200 | 10.24 | 6.88 | 54.3 | 54.7 |
Appendix figures & tables15 assets
Supplementary material from the paper’s appendix.
Appendix
| First-stage model | Cache features | VTC video targets |
|---|---|---|
| VideoPrism-LvT-B | VideoPrism-Base | VideoPrism-LvT-B |
| VideoCLIP-XL (ZS) | VideoCLIP-XL | VideoCLIP-XL |
| VideoCLIP-XL (FT) | VideoCLIP-XL | VideoCLIP-XL |
| PE-Core-B | PE-Core-B | PE-Core-B |
| CLIP ViT-B/16 (ZS) | CLIP ViT-B/16 | CLIP ViT-B/16 |
| Cache path | ||
|---|---|---|
| Host RAM | 5.79 | 7.83 |
| NVMe, warm | 6.3 | 7.8 |
| NVMe, cold | 7.4 | 11.1 |
| Representation | Visual payload per video |
|---|---|
| Uncompressed VideoPrism frame-and-patch representation | 6.02 MiB |
| VEDJE, 64 tokens | 48 KiB |
| VEDJE, 16 tokens | 12 KiB |
| Method | T2V R@1 |
|---|---|
| Released coarse checkpoint (reproduced) | 42.8 |
| EERCF published | 47.8 |
| VEDJE on released-checkpoint candidates | 49.1 |
| Variant | Joint tokens? | Learned prior? | T2V R@1 | V2T R@1 |
|---|---|---|---|---|
| Stage-1 baseline | ✗ | ✗ | 50.1 | 49.8 |
| Prior stream only | ✗ | ✓ | 50.1 | 49.8 |
| Joint stream only | ✓ | ✗ | 43.8 | 41.6 |
| Full VEDJE | ✓ | ✓ | 54.6 | 54.6 |
| Text-to-video | Video-to-text | |||||
| Dataset | Stage 1 | VEDJE | Gain | Stage 1 | VEDJE | Gain |
| VideoPrism first stage | ||||||
| MSR-VTT | 50.1 | 54.6 | +4.5 | 49.8 | 54.6 | +4.8 |
| MSVD | 55.6 | 56.7 | +1.1 | 83.4 | 83.7 | +0.3 |
| DiDeMo | 47.1 | 48.2 | +1.1 | 47.3 | 47.9 | +0.6 |
| ActivityNet | 48.8 | 50.6 | +1.8 | 47.9 | 48.3 | +0.4 |
| Dataset | Dir. | Stage-1 R@1 | Reranked R@1 | Gain | Reranked R@5 | Reranked R@10 |
| VideoPrism-LvT-B ( Zhao et al., 2024 ) stage 1 | ||||||
| MSR-VTT | T2V | 50.1 | 54.6 | +4.5 | 76.9 | 85.4 |
| MSR-VTT | V2T | 49.8 | 54.6 | +4.8 | 79.3 | 86.7 |
| MSVD | T2V | 55.6 | 56.7 | +1.1 | 83.7 | 89.8 |
| MSVD | V2T | 83.4 | 83.7 | +0.3 | 96.0 | 98.7 |
| DiDeMo | T2V | 47.1 | 48.2 | +1.1 | 72.9 | 80.3 |
| Precision | T2V R@1 | T2V R@5 | T2V R@10 | V2T R@1 | V2T R@5 | V2T R@10 |
|---|---|---|---|---|---|---|
| BF16 reference | 54.4 | 78.9 | 85.6 | 53.1 | 78.4 | 86.5 |
| FP8-E4M3 | 54.4 | 78.9 | 85.5 | 53.2 | 78.4 | 86.5 |
| FP4-E2M1 | 54.0 | 78.8 | 85.5 | 53.2 | 78.4 | 86.5 |
| Variant | T2V R@1 | T2V R@5 | T2V R@10 | V2T R@1 | V2T R@5 | V2T R@10 |
|---|---|---|---|---|---|---|
| 16 tokens w/o | 52.5 | 78.1 | 84.5 | 52.6 | 77.8 | 85.7 |
| 16 tokens with | 54.4 | 78.9 | 85.6 | 53.1 | 78.4 | 86.5 |
| Recall | Current-frame reconstruction | Absolute future feature | Future delta (default) |
|---|---|---|---|
| T2V R@1 | 52.8 | ||
| T2V R@5 | 78.2 | 76.9 | |
| T2V R@10 | 85.3 | 85.4 | |
| V2T R@1 | 53.8 | ||
| V2T R@5 | 78.6 | 79.3 | |
| V2T R@10 | 86.5 | 86.7 |
| Active training objectives | T2V R@1 | T2V R@5 | T2V R@10 | V2T R@1 | V2T R@5 | V2T R@10 |
|---|---|---|---|---|---|---|
| Stage-1 only ( Zhao et al., 2024 ) No reranker losses | 50.1 | – | – | 49.8 | – | – |
| 52.8 | 76.9 | 84.0 | 50.6 | 77.4 | 85.2 | |
| 52.7 | 76.2 | 84.2 | 51.4 | 76.0 | 84.9 | |
| 54.0 | 77.8 | 85.3 | 53.6 | 77.9 | 87.0 | |
| Full: | 54.6 | 76.9 | 85.4 | 54.6 | 79.3 | 86.7 |
| Horizon | T2V R@1 | T2V R@5 | T2V R@10 | V2T R@1 | V2T R@5 | V2T R@10 |
|---|---|---|---|---|---|---|
| 53.8 | 78.2 | 85.5 | 53.0 | 78.7 | 86.6 | |
| (default) | 54.6 | 76.9 | 85.4 | 54.6 | 79.3 | 86.7 |
| 54.4 | 78.7 | 85.6 | 53.0 | 78.2 | 86.9 |
| Variant | Added reranker params | T2V R@1 | T2V R@5 | T2V R@10 | V2T R@1 | V2T R@5 | V2T R@10 |
|---|---|---|---|---|---|---|---|
| Default VEDJE | 33M | 54.6 | 76.9 | 85.4 | 54.6 | 79.3 | 86.7 |
| BERT-base text encoder ( Devlin et al., 2019 ) | 110M | 52.9 | 77.5 | 85.5 | 54.0 | 77.9 | 86.5 |
| Method | Online parameters | R@1 (%) |
|---|---|---|
| Text encoding and learned scoring | ||
| CLIP4Clip B/32 | 63.4M | 44.5 |
| CLIP4Clip B/16 (reproduced) | 63.4M | 46.4 |
| X-CLIP B/16 | 64.0M | 49.3 |
| EERCF B/16 | 63.7M | 49.9 |
| Video-ColBERT SigLIP-B/16 | 110.3M | 51.5 |