Natural language descriptions can provide rich semantic representations of audio-visual urban scenes, yet datasets that jointly describe both auditory and visual information remain limited. In this paper, we introduce AVSD-Scenes, a paired audio-visual scene description dataset for urban environments. The dataset contains 12,291 audio-visual scene descriptions generated from the TAU Urban Audio-Visual Scenes dataset. To construct the dataset, we first generate audio- and visual-based descriptions using Qwen2-Audio-7B and Qwen2.5-VL-7B, respectively. These modality-specific descriptions are then combined using large language models, namely Qwen3-14B, Mistral-Small-3.2-24B-Instruct-2506, and Gemma-3-27B-it, to produce multimodal descriptions that capture complementary information from both modalities. We benchmark AVSD-Scenes using semantic alignment, cross-modal retrieval, scene classification, LLM-as-a-judge evaluation, and human subjective assessment. Results show that multimodal descriptions improve semantic alignment and cross-modal retrieval performance compared with modality-specific descriptions while preserving strong scene-discriminative information. The generated descriptions achieve up to 94.5% accuracy in urban scene classification, while combining audio, visual, and description embeddings further improves accuracy to 95.4%. Furthermore, the descriptions remain highly scene-discriminative even when scene labels are removed from the prompting instructions, indicating that they capture semantic information derived from the audio-visual content rather than merely reflecting label information.
Figures & tables
Figure 1: Comparison of AVSD-Scenes descriptions with existing descriptions of an audio-visual scene example “park-stockholm-246-7354” taken from TAU Urban Audio-Visual Scenes 2021 development set [ 15 ] .
Figure 2: Dataset construction pipeline. All the models are loaded using 4-bit NormalFloat quantization with double quantization and FP16 computation for efficient inference.
Description
CLIP / CLAP
IMAGEBIND
Text–Visual
Text–Audio
Avg.
Text–Visual
Text–Audio
Avg.
Modality-Specific
.3299
.2143
.2721
.3627
.0556
.2092
Multimodal (Qwen3)
.3187
.2854
.3021
.3425
.1696
.2561
Multimodal (Mistral)
.3018
.2613
.2816
.3165
.1615
.2390
Multimodal (Gemma)
.3083
.2735
.2909
.3218
.1599
.2409
Table 1: Cross-modal similarity of the generated descriptions. Avg. denotes the mean of the text–visual and text–audio similarity.
Description
CLAP / CLIP
IMAGEBIND
T → A
T → V
Avg.
T → A
T → V
Avg.
R@1
R@5
R@10
R@1
R@5
R@10
R@1
R@5
R@10
R@1
R@5
R@10
R@1
R@5
R@10
R@1
R@5
R@10
Modality-Specific
24.99
49.97
60.85
85.19
94.57
97.13
55.09
72.27
78.99
24.40
40.95
51.82
90.76
97.01
98.39
57.58
68.98
75.11
Multimodal (Qwen3)
41.43
68.00
77.06
84.49
93.98
96.70
62.96
80.99
86.88
58.55
82.12
89.87
90.56
96.89
98.23
74.56
89.51
94.05
Multimodal (Mistral)
37.00
63.82
73.20
85.47
94.91
97.35
61.23
79.36
85.27
59.78
82.52
90.23
89.57
96.72
98.23
74.68
89.62
94.23
Multimodal (Gemma)
38.35
65.74
75.40
83.04
93.50
96.36
60.69
79.62
85.88
56.46
78.66
86.05
87.78
95.17
97.25
72.12
86.92
91.65
Table 2: Scene-level cross-modal retrieval performance (%). Avg. denotes the mean of text–audio (T → A) and text–visual (T → V) retrieval.
Figure 3: SVM-based urban scene classification accuracy for different combinations of audio (A), visual (V), and multimodal description (D) obtained using Qwen3, Mistral, and Gemma models.
Figure 4: t-SNE visualizations of audio, visual, and Mistral-generated scene-description embeddings, and their multimodal combinations.
Figure 5: LLM-as-a-Judge evaluation of multimodal descriptions generated by different models. Scores are reported on a five point scale (higher is better). AF: audio fidelity; VF: visual fidelity; EC: event coverage; Halluc.: hallucination; CMC: cross-modal consistency; Gram.: grammar; Flue.: fluency.