TF-PRVR: Training-Free Partially Relevant Video Retrieval
Organizations: Chung-Ang University
Abstract
Partially Relevant Video Retrieval (PRVR) aims to retrieve untrimmed videos containing moments relevant to a given text query. Despite recent progress, existing PRVR methods suffer from two key limitations: a fixed video decomposition scheme that causes semantic dilution, and source-domain overfitting induced by task-specific training. In this paper, we propose TF-PRVR, the first training-free framework for PRVR. TF-PRVR leverages frozen vision-language features to construct video-specific hierarchical representations. It derives temporal semantic signals from frame-level features and applies frequency-based multi-scale analysis to identify adaptive temporal boundaries, producing hierarchical segments with coherent event-level semantics. Built on these segments, TF-PRVR constructs a unified multi-scale graph and propagates query relevance across temporally and semantically related nodes. A moment-aware scoring strategy then aggregates temporally aligned relevance across scales, emphasizing consistently supported moments while suppressing isolated false responses. Without task-specific training, TF-PRVR preserves the general-purpose alignment capability of pre-trained vision-language models and avoids dataset-specific overfitting. Extensive experiments demonstrate consistent performance across datasets with diverse visual and temporal characteristics, suggesting a practical direction for training-free PRVR.
Figures & tables
| Source Training Data | Method | R@1 | R@5 | R@10 | R@100 | SumR |
| Training-based methods with CLIP-L Radford et al. (2021) | ||||||
| Charades-STA Gao et al. (2017) | MS-SL Dong et al. (2022) | 0.4 | 2.2 | 3.9 | 17.3 | 23.8 |
| GMMFormer Wang et al. (2024b) | 0.4 | 2.6 | 4.3 | 18.8 | 26.1 | |
| GMMFormerv2 Wang et al. (2024a) | 0.6 | 2.8 | 4.6 | 19.6 | 27.6 | |
| MSC-PRVR Moon et al. (2025b) | 0.7 | 2.8 | 4.5 | 20.1 | 28.1 | |
| TVR Lei et al. (2020) | MS-SL Dong et al. (2022) | 6.9 | 19.8 | 26.5 | 59.1 | 112.3 |
| Method | ActivityNet Captions Krishna et al. (2017) | Charades-STA Gao et al. (2017) | TVR Lei et al. (2020) | ||||||||||||
| R@1 | R@5 | R@10 | R@100 | SumR | R@1 | R@5 | R@10 | R@100 | SumR | R@1 | R@5 | R@10 | R@100 | SumR | |
| Training-based methods with ResNet152 He et al. (2016) + I3D Carreira and Zisserman (2017) + RoBERTa Liu et al. (2019) | |||||||||||||||
| MS-SL Dong et al. (2022) | 7.1 | 22.5 | 34.7 | 75.8 | 140.1 | 1.8 | 7.1 | 11.8 | 47.7 | 68.4 | 13.5 | 32.1 | 43.4 | 83.4 | 172.4 |
| PEAN Jiang et al. (2023) | 7.4 | 23.0 | 35.5 | 75.9 | 141.8 | 2.7 | 8.1 | 13.5 | 50.3 | 74.7 | 13.5 | 32.8 | 44.1 | 83.9 | 174.2 |
| LH Fang et al. (2024) | 7.4 | 23.5 | 35.8 | 75.8 | 142.4 | 2.1 | 7.5 | 12.9 | 50.1 | 72.7 | 13.2 | 33.2 | 44.4 | 85.5 | 176.3 |
| BGM-Net Yin et al. (2024) | 7.2 | 23.8 | 36.0 | 76.9 | 143.9 | 1.9 | 7.4 | 12.2 | 50.1 | 71.6 | 14.1 | 34.7 | 45.9 | 85.2 | 179.9 |
| Dataset | Method | Source Training Data | R@1 | R@5 | R@10 | R@100 | SumR |
| QVHighlights Lei et al. (2021) | GMMFormerv2 Wang et al. (2024a) | Act+Cha+TVR | 18.6 | 26.5 | 29.8 | 50.6 | 125.5 |
| MSC-PRVR Moon et al. (2025b) | Act+Cha+TVR | 20.1 | 29.2 | 33.0 | 49.8 | 132.1 | |
| CLIP Zero-shot Radford et al. (2021) | None | 21.0 | 41.4 | 52.3 | 82.9 | 197.6 | |
| TF-PRVR (Ours) | None | 23.9 | 47.5 | 57.6 | 89.5 | 218.5 | |
| EPIC-Kitchens-100 Damen et al. (2022) | GMMFormerv2 Wang et al. (2024a) | Act+Cha+TVR | 2.5 | 8.1 | 12.4 | 21.1 | 44.1 |
| MSC-PRVR Moon et al. (2025b) | Act+Cha+TVR | 2.8 | 8.5 | 13.1 | 22.9 | 47.3 |
| Source Training Data | Method | R@1 | R@5 | R@10 | R@100 | SumR |
| Training-based methods with CLIP-L Radford et al. (2021) | ||||||
| Charades-STA Gao et al. (2017) | MS-SL Dong et al. (2022) | 1.8 | 7.1 | 11.8 | 47.7 | 68.4 |
| GMMFormer Wang et al. (2024b) | 2.1 | 7.8 | 12.5 | 50.5 | 72.9 | |
| GMMFormerv2 Wang et al. (2024a) | 2.1 | 7.8 | 12.5 | 50.5 | 72.9 | |
| MSC-PRVR Moon et al. (2025b) | 3.2 | 12.6 | 20.1 | 63.8 | 99.7 | |
| ActivityNet Captions Krishna et al. (2017) | MS-SL Dong et al. (2022) | 0.8 | 3.4 | 6.2 | 31.7 | 42.1 |
| Source Training Data | Method | R@1 | R@5 | R@10 | R@100 | SumR |
| Training-based methods with CLIP-L Radford et al. (2021) | ||||||
| TVR Lei et al. (2020) | MS-SL Dong et al. (2022) | 31.9 | 57.6 | 67.7 | 93.8 | 251.0 |
| GMMFormer Wang et al. (2024b) | 29.8 | 54.2 | 64.6 | 92.5 | 241.1 | |
| GMMFormerv2 Wang et al. (2024a) | 34.0 | 59.7 | 69.8 | 94.6 | 258.1 | |
| MSC-PRVR Moon et al. (2025b) | 35.1 | 61.6 | 71.5 | 94.9 | 263.1 | |
| ActivityNet Captions Krishna et al. (2017) | MS-SL Dong et al. (2022) | 3.8 | 9.8 | 14.2 | 40.7 | 68.5 |
| Dataset | Segmentation | Mean Best tIoU | R@0.5 | GT Coverage | Fragmentation | Irrelevant Inclusion | |
| ActivityNet Captions Krishna et al. (2017) | Fixed scheme | 0.177 | 7.2% | 0.187 | 12.41 | 0.082 | 0.0864 |
| Ours | 0.315 | 24.8% | 0.354 | 5.86 | 0.041 | 0.0592 | |
| Charades-STA Gao et al. (2017) | Fixed scheme | 0.139 | 0.1% | 0.151 | 9.78 | 0.057 | 0.0109 |
| Ours | 0.278 | 14.2% | 0.291 | 4.85 | 0.031 | 0.0076 | |
| TVR Lei et al. (2020) | Fixed scheme | 0.348 | 27.0% | 0.367 | 5.37 | 0.076 | 0.0117 |
| Ours | 0.462 | 45.3% | 0.491 | 2.84 | 0.038 | 0.0082 |
| Boundary Strategy | Budget | Act Krishna et al. (2017) | Cha Gao et al. (2017) | TVR Lei et al. (2020) | QVH Lei et al. (2021) |
| Uniform fixed intervals (conventional) | 160 | 191.8 | 80.1 | 175.4 | 206.1 |
| Adjacent feature dissimilarity | 160 | 196.2 | 82.3 | 179.4 | 212.6 |
| Self-similarity matrix | 160 | 197.0 | 83.8 | 180.1 | 213.9 |
| Velocity-direction change (Ours) | 160 | 200.7 | 85.1 | 183.0 | 218.4 |
| Ours, unconstrained (video-specific) | Adaptive | 200.8 | 85.4 | 183.6 | 218.5 |
| Method | Segmentation | GP | MAS | Act | Cha | TVR | QVH |
| Zero-shot | – | – | – | 180.4 | 77.5 | 172.2 | 197.6 |
| + Single-scale mean pooling (32 clips) | Fixed | – | – | 183.5 | 78.1 | 173.2 | 199.8 |
| + Two-branch (clip + frame) | Fixed | – | – | 186.6 | 78.8 | 174.4 | 202.1 |
| + Multi-scale temporal pooling (3 scales) | Fixed | – | – | 188.2 | 79.2 | 175.0 | 203.5 |
| + Sliding-window multi-scale (exhaustive) | Fixed | – | – | 189.5 | 79.5 | 175.3 | 204.8 |
| + Graph propagation only | Fixed | – | 190.6 | 80.0 | 175.8 | 205.9 |
| Method | ActivityNet Captions Krishna et al. (2017) | Charades-STA Gao et al. (2017) | TVR Lei et al. (2020) | ||||||
| Full | Short | Long | Full | Short | Long | Full | Short | Long | |
| TF-PRVR (Ours) | 193.9 | 193.7 | 194.0 | 80.8 | 80.7 | 80.8 | 178.3 | 178.4 | 178.1 |
| Method | Indexed Query Latency (ms) | End-to-End Time (s) | Indexed Query Memory (GB) | Peak Deployment Memory (GB) |
| MS-SL Dong et al. (2022) | 54.33 | 48.74 | 7.57 | 8.29 |
| GMMFormerv2 Wang et al. (2024a) | 9.74 | 48.31 | 1.94 | 2.66 |
| MSC-PRVR Moon et al. (2025b) | 9.78 | 48.36 | 0.48 | 1.20 |
| TF-PRVR (Ours) | 17.24 | 49.08 | 3.80 | 4.52 |