Partially Relevant Video Retrieval (PRVR) seeks to retrieve untrim-med videos containing a moment that matches a text query, without temporal annotations. The relevant moment may last only seconds within a video spanning several minutes, creating an extremely low signal-to-noise ratio that makes PRVR more challenging than standard full-video retrieval. This task presents two intertwined challenges: (1) signal dilution, where coarse global representations blur the brief relevant signal into the dominant irrelevant surroundings;(2) curvature rigidity, where embedding all videos in the same fixed-geometry space distorts representations for videos that range from flat atomic events to deep compositional hierarchies. Existing PRVR methods have improved moment selection and cross-modal matching, but they still typically encode all videos in a single fixed-curvature retrieval space, limiting their ability to model diverse video structures. To address both challenges, we propose CurvSpec, a framework that learns content-adaptive curvature for video retrieval representations rather than imposing a fixed geometric prior. CurvSpec processes features through parallel Euclidean and hyperbolic attention layers, with independently learned curvatures assigned to the hyperbolic layers, and a content-aware fusion mechanism routes each input to its most suitable geometric regime. To further suppress signal dilution, CurvSpec represents each video with semantic centroids whose number is determined by the video's content complexity, projects them onto the learned manifold, and matches each query against its nearest centroid by geodesic distance. Experiments on ActivityNet Captions, TVR, and Charades-STA demonstrate state-of-the-art retrieval performance.
Figures & tables
Figure 1 . Motivation and key ideas. (a) Videos exhibit diverse structures (Video 1 to 3, from flat to deep hierarchies). Neither Euclidean space nor a single fixed hyperbolic curvature accommodates all of them; CurvSpec routes each video to its optimal geometry. (b) Global pooling dilutes relevant signals into an ambiguous single vector; semantic centroids decompose the video into coherent regions, allowing the query to match the nearest relevant centroid free from irrelevant interference.
Figure 2 . Overview of CurvSpec. (a) Text queries and video frames/clips are encoded and processed by the CurvSpec block to yield enriched representations Zc and Zf . Clip features are further decomposed into semantic centroids, which, together with the query, are projected onto the Lorentz manifold for fine-grained partial relevance matching. (b) The CurvSpec block processes input through parallel Euclidean and Lorentz attention branches, each at a distinct geometry. A content-aware fusion mechanism adaptively combines the multi-geometry representations.
Method
ActivityNet Captions
TVR
Charades-STA
R@1
R@5
R@10
R@100
SumR
R@1
R@5
R@10
R@100
SumR
R@1
R@5
R@10
R@100
SumR
Text-to-Video Retrieval
HGR ( Chen et al., 2020 )
4.0
15.0
24.8
63.2
107.0
1.7
4.9
8.3
35.2
50.1
1.2
3.8
7.3
33.4
45.7
DE++ ( Dong et al., 2022b )
5.3
18.4
29.2
68.0
121.0
8.8
21.9
30.2
67.4
128.3
1.7
5.6
9.6
37.1
54.1
RIVRL ( Dong et al., 2022c )
5.2
18.0
28.2
66.4
117.8
9.4
23.4
32.2
70.6
135.6
1.6
5.6
9.4
37.7
54.3
CLIP4Clip ( Luo et al., 2022 )
5.9
19.3
30.4
71.6
127.3
9.9
24.3
34.3
72.5
141.0
1.8
6.5
10.9
44.2
63.4
Table 1. Retrieval performance on ActivityNet Captions, TVR, and Charades-STA. Bold and underline denote the best and second-best results, respectively. “–” indicates unavailable results.
1The Hong Kong University of Science and Technology (Guangzhou) · University of Electronic Science and Technology of China · 3The Hong Kong University of Science and Technology
Tsinghua Shenzhen International Graduate School, Tsinghua University, Shenzhen, China · Peng Cheng Laboratory, Shenzhen, China · Harbin Institute of Technology, Shenzhen, China