Unlocking Spatial Grounding in Large Audio-Visual Retrieval models
Authors: Hugo Malard, Michel Olvera, Sanjeel Parekh, Gaël Richard, Slim Essid, Stéphane Lathuilière
Organizations: LTCI, Télécom Paris, Institut Polytechnique de Paris, France · Meta, Reality Labs Research · NVIDIA, France · Inria at Université Grenoble Alpes, CNRS, LJK, France
Weak supervision sets a practical regime for audio-visual sound source localization as dense spatial annotations are costly to obtain at scale. The task, however, remains challenging, as models must locate sound sources from temporally aligned audio-visual data without pixel-level supervision. Recent large-scale audio-visual retrieval models, trained at unprecedented scale, encode rich multimodal structure. We show their latent representations, though optimized for global alignment, can nonetheless enable fine-grained spatial grounding. While spatial detail is progressively lost in the upper layers of retrieval backbones due to global pooling, intermediate visual tokens retain highly structured spatial information. To exploit this, we introduce LAIP (\emph{Localization via Audio-Informed Pooling}), a framework that employs a lightweight \emph{Audio-informed Spatial Pooling} (AiSP) to replace the standard global aggregation module. By querying intermediate visual tokens with audio aligned at the frame level, LAIP recovers localized spatial information that is otherwise discarded by the retrieval pipeline, with the largest gains observed for PE-AV, a stack with underlying temporal aggregation. Our approach achieves state-of-the-art performance on AVSBench and AVATAR, nearly doubling previous results on the latter, improving average CIoU from 13.21 to 26.22.
Figures & tables
Figure 1: Global alignment suppresses spatial information in frame and video-level features. Earlier frame-encoder layers retain locality but are incompatible with the global multimodal space. LAIP learns locally grounded representations compatible with the audio space.
Figure 2: LAIP: Localization via Audio-Informed Pooling . The audio tokens are first forwarded to a local aggregator to match visual frame rate, and then used to pool the frame representation through our proposed localization module . The pooled representations are used as input to the video encoder , which outputs the final representation used to compute the sigmoid loss.
Figure 3: Overview of our localization and AiSP modules, with K=2 . Starting from PE visual tokens extracted from the image branch at the 16 -th frame-encoder layer, a frame-aligned audio token progressively pools the spatial grid into one sound-conditioned visual token per frame.
(1) Single-sound
(2) Mixed-sound
(3) Multi-entity
(4) Off-screen
Method
CIoU(%) ↑
AUC(%) ↑
CIoU(%) ↑
AUC(%) ↑
CIoU(%) ↑
AUC(%) ↑
TN(%) ↑
EZ-VSL(10k) ( Mo and Morgado, 2022b )
9.66
11.07
8.16
9.35
6.87
8.32
96.91
EZ-VSL(144k) ( Mo and Morgado, 2022b )
10.92
12.22
6.97
8.34
5.80
7.42
96.47
FNAC (144k) ( Sun et al., 2023 )
11.17
13.21
5.87
7.73
3.77
5.79
92.84
FNAC (10k) ( Sun et al., 2023 )
13.00
14.19
10.99
12.08
10.31
11.45
87.52
TAVLO(10k) ( Choi et al., 2025 )
13.42
14.08
14.13
14.52
12.08
12.69
91.18
Table 1: Segmentation performance comparison across AVATAR scenarios.
Method
S4
MS3
mask-IoU ↑
F-score ↑
mask-IoU ↑
F-score ↑
SLAVC ( Mo and Morgado, 2022a )
28.10
34.60
24.37
25.56
MarginNCE ( Park et al., 2023 )
33.27
45.33
27.31
31.56
FNAC ( Sun et al., 2023 )
27.15
31.40
21.98
22.50
Alignment ( Senocak et al., 2023 )
29.60
35.90
-
-
TACO ( Malard et al., 2026 )
29.68
41.91
25.88
30.72
Table 2: Quantitative results on the AVSBench test sets.
Method
Retrieval model
m-IoU ↑
mAP ↑
DAVENet ( Hsu et al., 2019 )
✗
17.0
16.8
DenseAV ( Hamilton et al., 2024 )
✗
25.5
32.4
TACO ( Malard et al., 2026 )
✗
27.74
35.75
ImageBind ( Girdhar et al., 2023 )
✓
18.3
18.1
CAVMAE ( Gong et al., 2023 )
✓
20.6
21.2
CAV-MAE Sync ( Araujo et al., 2025 )
✓
22.7
22.6
Table 3: Quantitative results on the ADE-SP dataset.
Method
CIoU(%) ↑
AUC(%) ↑
LAIP (Ours)
26.22
26.34
μ=0
23.01
23.27
No null token
23.11
23.25
μ=λ=0
19.45
19.53
λ=0
17.52
17.57
Table 4: AVATAR ablations (averaged over scenarios): regularization (left) and baselines (right).
Figure 4: TACO–LAIP comparison on S4. LAIP produces smoother maps with fewer outliers.
Appendix figures & tables14 assets
Supplementary material from the paper’s appendix.
Appendix
(1) Single-sound
(2) Mixed-sound
(3) Multi-entity
(4) Off-screen
Method
CIoU(%) ↑
AUC(%) ↑
CIoU(%) ↑
AUC(%) ↑
CIoU(%) ↑
AUC(%) ↑
TN(%) ↑
LAIP (Ours)
27.63
27.77
27.35
27.40
23.69
23.85
90.94
μ=0
24.46
24.75
24.60
24.72
19.97
20.34
92.45
No null token
24.37
24.55
24.18
24.20
20.79
21.01
90.89
μ=λ=0
19.59
19.70
21.18
21.18
17.57
17.71
90.76
λ=0
17.13
17.17
19.82
19.85
15.61
15.69
89.35
Appendix
Table 6: Ablation study on AVATAR across the four evaluation scenarios.
(1) Single-sound
(2) Mixed-sound
(3) Multi-entity
(4) Off-screen
Method
CIoU(%) ↑
AUC(%) ↑
CIoU(%) ↑
AUC(%) ↑
CIoU(%) ↑
AUC(%) ↑
TN(%) ↑
LAIP (Ours)
27.63
27.77
27.35
27.40
23.69
23.85
90.94
LAIP with global audio query
26.40
26.52
27.91
27.97
22.57
22.69
91.27
Gradient-based
4.35
5.06
3.49
4.39
4.58
5.22
88.17
LAIP on last PE layer
16.45
17.02
17.17
17.42
14.14
14.70
89.76
LAIP video encoder scratch
19.03
19.35
22.80
22.88
17.59
17.82
92.03
Appendix
Table 7: Ablation study on AVATAR across the four evaluation scenarios.
(1) Single-sound
(2) Mixed-sound
(3) Multi-entity
(4) Off-screen
Method
CIoU(%) ↑
AUC(%) ↑
CIoU(%) ↑
AUC(%) ↑
CIoU(%) ↑
AUC(%) ↑
TN(%) ↑
Ours
27.63
27.77
27.35
27.40
23.69
23.85
90.94
LAIP Pooler inverted bottlneck (4 3 2)
13.67
13.94
15.01
15.12
11.95
12.38
90.83
LAIP two poolers
21.33
21.45
23.21
23.26
17.51
17.70
90.52
Appendix
Table 8: Ablation study on AVATAR across the four evaluation scenarios.
Figure 6: Impact of the thresholding on attention maps.
Figure 7: Audio–visual similarity of hard-negative pairs as a function of the attention assigned to the null token.
Vision encoder
CIoU
AUC
PE-AV tokens
7.05
7.54
ResNet-18
7.78
9.03
Appendix
Table 9: Performance of TAVLO adapted to PE tokens
(1) Single-sound
(2) Mixed-sound
(3) Multi-entity
(4) Off-screen
Method
CIoU(%) ↑
AUC(%) ↑
CIoU(%) ↑
AUC(%) ↑
CIoU(%) ↑
AUC(%) ↑
TN(%) ↑
LAIP (PE-AV L16)
27.63
27.77
27.35
27.40
23.69
23.85
90.94
LAIP (PE-AV L16, video encoder scratch)
19.03
19.35
22.80
22.88
17.59
17.82
92.03
LAIP (ImageBind L24)
18.29
18.56
16.14
16.30
15.60
16.02
90.05
LAIP (ImageBind L32)
5.83
6.88
5.76
6.79
5.14
6.28
89.93
Appendix
Table 10: LAIP performance with PE-AV and ImageBind features on the four AVATAR scenarios. Both ImageBind variants use a randomly initialized downstream video encoder.
(1) Single-sound
(2) Mixed-sound
(3) Multi-entity
(4) Off-screen
PE-AV layer
CIoU(%) ↑
AUC(%) ↑
CIoU(%) ↑
AUC(%) ↑
CIoU(%) ↑
AUC(%) ↑
TN(%) ↑
8
19.51
19.68
20.32
20.36
16.09
16.33
89.73
14
25.56
25.62
22.77
22.79
20.70
20.84
90.59
16
27.63
27.77
27.35
27.40
23.69
23.85
90.94
18
23.18
23.27
23.87
23.89
20.62
20.73
90.78
24
16.45
17.02
17.17
17.42
14.14
14.70
89.76
Appendix
Table 11: Layer-wise LAIP performance on the four AVATAR scenarios using PE-AV frame-encoder features.
(1) Single-sound
(2) Mixed-sound
(3) Multi-entity
(4) Off-screen
Configuration
CIoU(%) ↑
AUC(%) ↑
CIoU(%) ↑
AUC(%) ↑
CIoU(%) ↑
AUC(%) ↑
TN(%) ↑
LAIP (default)
27.63
27.77
27.35
27.40
23.69
23.85
90.94
μ=0
24.46
24.75
24.60
24.72
19.97
20.34
92.45
μ=λ=0
19.59
19.70
21.18
21.18
17.57
17.71
90.76
λ=0
17.13
17.17
19.82
19.85
15.61
15.69
89.35
μ=0.001
24.92
25.10
24.83
24.87
20.56
20.83
91.32
Appendix
Table 12: Hyperparameter sensitivity on the four AVATAR scenarios. Each row changes only the indicated setting relative to the default LAIP configuration.
Resource
PE-AV
PE-AV + LAIP
Overhead
Compute
6.67 TFLOPs
6.95 TFLOPs
+4.22%
Peak memory
8.45 GiB
8.79 GiB
+4.0%
Appendix
Table 13: Incremental computational overhead of LAIP for a 10-second video-audio pair.
Method
Vision encoding
Audio encoding
Localizer
Total
Localizer/total
DenseAV
27.0
6.72
0.12‡
33.9
0.35%
TACO
55.0
0.63
11.1†
66.8
16.7%
CAV-MAE Sync
29.2
30.8
0.00024
60.0
0.0004%
LAIP (224 px)
91.9
52.7
6.3
150.9
4.2%
Appendix
Table 14: MAC breakdown ( ×1010 ) of the 16-frame totals above. † Estimated decomposition. ‡ Estimated similarity map.
Figure 8: PE-AV does not encode audio-visual spatial correspondences. The norm of the gradient of the input and output tokens of the frame encoder with respect to the audio visual similarity does not exhibit interpretable patterns. The attention of the original attention pooling of PE does not either. However our model focuses exactly on the sounding regions of the image.
Figure 9: Qualitative samples of LAIP on the MS3 (first row) an S4 (second raw) dataset.