Unlocking Spatial Grounding in Large Audio-Visual Retrieval models
Authors: Hugo Malard, Michel Olvera, Sanjeel Parekh, Gaël Richard, Slim Essid, Stéphane Lathuilière
Organizations: LTCI, Télécom Paris, Institut Polytechnique de Paris, France · Meta, Reality Labs Research · NVIDIA, France · Inria at Université Grenoble Alpes, CNRS, LJK, France
Weak supervision sets a practical regime for audio-visual sound source localization as dense spatial annotations are costly to obtain at scale. The task, however, remains challenging, as models must locate sound sources from temporally aligned audio-visual data without pixel-level supervision. Recent large-scale audio-visual retrieval models, trained at unprecedented scale, encode rich multimodal structure. We show their latent representations, though optimized for global alignment, can nonetheless enable fine-grained spatial grounding. While spatial detail is progressively lost in the upper layers of retrieval backbones due to global pooling, intermediate visual tokens retain highly structured spatial information. To exploit this, we introduce LAIP (\emph{Localization via Audio-Informed Pooling}), a framework that employs a lightweight \emph{Audio-informed Spatial Pooling} (AiSP) to replace the standard global aggregation module. By querying intermediate visual tokens with audio aligned at the frame level, LAIP recovers localized spatial information that is otherwise discarded by the retrieval pipeline, with the largest gains observed for PE-AV, a stack with underlying temporal aggregation. Our approach achieves state-of-the-art performance on AVSBench and AVATAR, nearly doubling previous results on the latter, improving average CIoU from 13.21 to 26.22.
Figures & tables
Figure 1: Global alignment suppresses spatial information in frame and video-level features. Earlier frame-encoder layers retain locality but are incompatible with the global multimodal space. LAIP learns locally grounded representations compatible with the audio space.
Figure 2: LAIP: Localization via Audio-Informed Pooling . The audio tokens are first forwarded to a local aggregator to match visual frame rate, and then used to pool the frame representation through our proposed localization module . The pooled representations are used as input to the video encoder , which outputs the final representation used to compute the sigmoid loss.
Figure 3: Overview of our localization and AiSP modules, with K=2 . Starting from PE visual tokens extracted from the image branch at the 16 -th frame-encoder layer, a frame-aligned audio token progressively pools the spatial grid into one sound-conditioned visual token per frame.
(1) Single-sound
(2) Mixed-sound
(3) Multi-entity
(4) Off-screen
Method
CIoU(%) ↑
AUC(%) ↑
CIoU(%) ↑
AUC(%) ↑
CIoU(%) ↑
AUC(%) ↑
TN(%) ↑
EZ-VSL(10k) ( Mo and Morgado, 2022b )
9.66
11.07
8.16
9.35
6.87
8.32
96.91
EZ-VSL(144k) ( Mo and Morgado, 2022b )
10.92
12.22
6.97
8.34
5.80
7.42
96.47
FNAC (144k) ( Sun et al., 2023 )
11.17
13.21
5.87
7.73
3.77
5.79
92.84
FNAC (10k) ( Sun et al., 2023 )
13.00
14.19
10.99
12.08
10.31
11.45
87.52
TAVLO(10k) ( Choi et al., 2025 )
13.42
14.08
14.13
14.52
12.08
12.69
91.18
Table 1: Segmentation performance comparison across AVATAR scenarios.
Method
S4
MS3
mask-IoU ↑
F-score ↑
mask-IoU ↑
F-score ↑
SLAVC ( Mo and Morgado, 2022a )
28.10
34.60
24.37
25.56
MarginNCE ( Park et al., 2023 )
33.27
45.33
27.31
31.56
FNAC ( Sun et al., 2023 )
27.15
31.40
21.98
22.50
Alignment ( Senocak et al., 2023 )
29.60
35.90
-
-
TACO ( Malard et al., 2026 )
29.68
41.91
25.88
30.72
Table 2: Quantitative results on the AVSBench test sets.
Method
Retrieval model
m-IoU ↑
mAP ↑
DAVENet ( Hsu et al., 2019 )
✗
17.0
16.8
DenseAV ( Hamilton et al., 2024 )
✗
25.5
32.4
TACO ( Malard et al., 2026 )
✗
27.74
35.75
ImageBind ( Girdhar et al., 2023 )
✓
18.3
18.1
CAVMAE ( Gong et al., 2023 )
✓
20.6
21.2
CAV-MAE Sync ( Araujo et al., 2025 )
✓
22.7
22.6
Table 3: Quantitative results on the ADE-SP dataset.
Method
CIoU(%) ↑
AUC(%) ↑
LAIP (Ours)
26.22
26.34
μ=0
23.01
23.27
No null token
23.11
23.25
μ=λ=0
19.45
19.53
λ=0
17.52
17.57
Table 4: AVATAR ablations (averaged over scenarios): regularization (left) and baselines (right).
Figure 4: TACO–LAIP comparison on S4. LAIP produces smoother maps with fewer outliers.
Appendix figures & tables14 assets
Supplementary material from the paper’s appendix.
Appendix
(1) Single-sound
(2) Mixed-sound
(3) Multi-entity
(4) Off-screen
Method
CIoU(%) ↑
AUC(%) ↑
CIoU(%) ↑
AUC(%) ↑
CIoU(%) ↑
AUC(%) ↑
TN(%) ↑
LAIP (Ours)
27.63
27.77
27.35
27.40
23.69
23.85
90.94
μ=0
24.46
24.75
24.60
24.72
19.97
20.34
92.45
No null token
24.37
24.55
24.18
24.20
20.79
21.01
90.89
μ=λ=0
19.59
19.70
21.18
21.18
17.57
17.71
90.76
λ=0
17.13
17.17
19.82
19.85
15.61
15.69
89.35
Appendix
Table 6: Ablation study on AVATAR across the four evaluation scenarios.
(1) Single-sound
(2) Mixed-sound
(3) Multi-entity
(4) Off-screen
Method
CIoU(%) ↑
AUC(%) ↑
CIoU(%) ↑
AUC(%) ↑
CIoU(%) ↑
AUC(%) ↑
TN(%) ↑
LAIP (Ours)
27.63
27.77
27.35
27.40
23.69
23.85
90.94
LAIP with global audio query
26.40
26.52
27.91
27.97
22.57
22.69
91.27
Gradient-based
4.35
5.06
3.49
4.39
4.58
5.22
88.17
LAIP on last PE layer
16.45
17.02
17.17
17.42
14.14
14.70
89.76
LAIP video encoder scratch
19.03
19.35
22.80
22.88
17.59
17.82
92.03
Appendix
Table 7: Ablation study on AVATAR across the four evaluation scenarios.
(1) Single-sound
(2) Mixed-sound
(3) Multi-entity
(4) Off-screen
Method
CIoU(%) ↑
AUC(%) ↑
CIoU(%) ↑
AUC(%) ↑
CIoU(%) ↑
AUC(%) ↑
TN(%) ↑
Ours
27.63
27.77
27.35
27.40
23.69
23.85
90.94
LAIP Pooler inverted bottlneck (4 3 2)
13.67
13.94
15.01
15.12
11.95
12.38
90.83
LAIP two poolers
21.33
21.45
23.21
23.26
17.51
17.70
90.52
Appendix
Table 8: Ablation study on AVATAR across the four evaluation scenarios.
Figure 6: Impact of the thresholding on attention maps.
Figure 7: Audio–visual similarity of hard-negative pairs as a function of the attention assigned to the null token.
Vision encoder
CIoU
AUC
PE-AV tokens
7.05
7.54
ResNet-18
7.78
9.03
Appendix
Table 9: Performance of TAVLO adapted to PE tokens
(1) Single-sound
(2) Mixed-sound
(3) Multi-entity
(4) Off-screen
Method
CIoU(%) ↑
AUC(%) ↑
CIoU(%) ↑
AUC(%) ↑
CIoU(%) ↑
AUC(%) ↑
TN(%) ↑
LAIP (PE-AV L16)
27.63
27.77
27.35
27.40
23.69
23.85
90.94
LAIP (PE-AV L16, video encoder scratch)
19.03
19.35
22.80
22.88
17.59
17.82
92.03
LAIP (ImageBind L24)
18.29
18.56
16.14
16.30
15.60
16.02
90.05
LAIP (ImageBind L32)
5.83
6.88
5.76
6.79
5.14
6.28
89.93
Appendix
Table 10: LAIP performance with PE-AV and ImageBind features on the four AVATAR scenarios. Both ImageBind variants use a randomly initialized downstream video encoder.
(1) Single-sound
(2) Mixed-sound
(3) Multi-entity
(4) Off-screen
PE-AV layer
CIoU(%) ↑
AUC(%) ↑
CIoU(%) ↑
AUC(%) ↑
CIoU(%) ↑
AUC(%) ↑
TN(%) ↑
8
19.51
19.68
20.32
20.36
16.09
16.33
89.73
14
25.56
25.62
22.77
22.79
20.70
20.84
90.59
16
27.63
27.77
27.35
27.40
23.69
23.85
90.94
18
23.18
23.27
23.87
23.89
20.62
20.73
90.78
24
16.45
17.02
17.17
17.42
14.14
14.70
89.76
Appendix
Table 11: Layer-wise LAIP performance on the four AVATAR scenarios using PE-AV frame-encoder features.
(1) Single-sound
(2) Mixed-sound
(3) Multi-entity
(4) Off-screen
Configuration
CIoU(%) ↑
AUC(%) ↑
CIoU(%) ↑
AUC(%) ↑
CIoU(%) ↑
AUC(%) ↑
TN(%) ↑
LAIP (default)
27.63
27.77
27.35
27.40
23.69
23.85
90.94
μ=0
24.46
24.75
24.60
24.72
19.97
20.34
92.45
μ=λ=0
19.59
19.70
21.18
21.18
17.57
17.71
90.76
λ=0
17.13
17.17
19.82
19.85
15.61
15.69
89.35
μ=0.001
24.92
25.10
24.83
24.87
20.56
20.83
91.32
Appendix
Table 12: Hyperparameter sensitivity on the four AVATAR scenarios. Each row changes only the indicated setting relative to the default LAIP configuration.
Resource
PE-AV
PE-AV + LAIP
Overhead
Compute
6.67 TFLOPs
6.95 TFLOPs
+4.22%
Peak memory
8.45 GiB
8.79 GiB
+4.0%
Appendix
Table 13: Incremental computational overhead of LAIP for a 10-second video-audio pair.
Method
Vision encoding
Audio encoding
Localizer
Total
Localizer/total
DenseAV
27.0
6.72
0.12‡
33.9
0.35%
TACO
55.0
0.63
11.1†
66.8
16.7%
CAV-MAE Sync
29.2
30.8
0.00024
60.0
0.0004%
LAIP (224 px)
91.9
52.7
6.3
150.9
4.2%
Appendix
Table 14: MAC breakdown ( ×1010 ) of the 16-frame totals above. † Estimated decomposition. ‡ Estimated similarity map.
Figure 8: PE-AV does not encode audio-visual spatial correspondences. The norm of the gradient of the input and output tokens of the frame encoder with respect to the audio visual similarity does not exhibit interpretable patterns. The attention of the original attention pooling of PE does not either. However our model focuses exactly on the sounding regions of the image.
Figure 9: Qualitative samples of LAIP on the MS3 (first row) an S4 (second raw) dataset.
Hour-scale audio-visual question answering is constrained by a context dilemma: dense whole-recording encoding rapidly exhausts context limits, whereas uniform temporal compression severely dilutes fine-grained acoustic and visual evidence. We introduce LEAP, a framework where the model retrieves its own evidence without placing the whole recording in one context. LEAP divides a recording into fixed-duration blocks, applying a lightweight localization pass to each block to score short candidate windows. The highest-ranked windows are pooled and re-encoded in a single bounded answer pass. Consequently, the answer input and peak context remain independent of the recording duration. By decoupling evidence localization from reasoning, our framework can localize candidate temporal windows over pre-computed transcripts without decoding media frames, while preserving fine-grained visual and non-speech evidence by routing the final answering pass over raw audio-visual streams. LEAP trains both stages: a localization LoRA improves the selected windows, and an answer LoRA improves the answers read from the same windows. The block grid natively supports causal queries, enabling LEAP to support streaming inference without streaming-specific training. Across several AVQA benchmarks, LEAP improves over the Qwen3-Omni-30B-A3B baseline by 4.5-16.8%, and transfers to a second omni-modal backbone, MiniCPM-o 4.5, surpassing its published results by 3.1-13.0%.
Understanding dynamic sound sources requires jointly determining what produces a sound, where the source is located, and how it moves over time. Yet existing audio-language models often represent clips as global acoustic events, while vision-language models lack the spatial audio cues needed to localize and track individual sources. To evaluate this missing capability, we introduce ST-OmniQA, a spatio-temporal audio-visual question-answering benchmark built from panoramic videos paired with synchronized first-order Ambisonics (FOA) audio of moving sound sources. It contains 40K videos and 400K question-answer pairs organized into four capability levels covering sound-event recognition, direction of arrival, source distance, motion trajectories, and temporally grounded audio-visual reasoning. Building on this benchmark, we propose ST-Omni-R1, which integrates FOA-derived semantic and trajectory representations with panoramic visual context and is trained through progressive curriculum learning and reasoning-tree reinforcement learning. ST-Omni-R1 achieves 77.83% average semantic accuracy across the four levels, compared with 37.28% for the best evaluated baseline. Results on three public spatial-audio benchmarks further indicate that its learned spatial and motion representations transfer beyond ST-OmniQA.
Zhi Zeng, Cheng Zhang, Zesheng Yang +9
Xi’an Jiaotong University · The Hong Kong Polytechnic University · National University of Singapore +2
Sound events are entities with semantic identities, locations, and trajectories, but current audio-language models usually reason about clips as global event content. Conversely, sound event localization models track source directions over time but offer limited semantic coverage for language reasoning. To address this gap, we introduce ST-AudioQA, a spatio-temporal audio QA dataset and benchmark built from first-order ambisonic (FOA) renderings of static and moving sound sources. Each scene provides source identity, activity, direction, distance, and motion metadata, enabling dense trajectory supervision and questions about what is sounding, where it is, how it moves, and how sources relate. We further propose ST-Audio Encoder, a time-resolved FOA audio encoder that learns event semantics together with source trajectories, and ST-AudioLM, which connects the audio tokens from the encoder to an LLM for spatio-temporal audio QA. Experiments show that this representation improves the semantic-localization tradeoff and yields stronger reasoning performance than static spatial and localization-oriented baselines.