We present a method for learning binaural sound localization using egomotion as a supervisory signal. Over the course of a video, the cameras direction to a sound source will change as the camera moves. We train an audio model to predict sound directions that are consistent with visual estimates of camera motion, which we obtain using traditional methods from multi-view geometry. This provides a weak but plentiful form of supervision that we combine with traditional binaural cues. To evaluate this method, we propose a dataset of real-world audio-visual videos with egomotion. We show that our model can successfully learn from real-world data and that it performs well on sound localization tasks
Figures & tables
Figure 1 : Supervising sound localization using egomotion from natural video. We use camera motion as a supervisory signal for stereo sound localization. Our model learns to predict changes in sound direction that correlate with changes in visually indicated camera motion. In contrast to prior self-supervised sound localization methods, our model is trained solely on in-the-wild natural videos, which often contain complex camera motions and noisy mixtures of different component sounds.
Figure 2 : The StereoWalks dataset. Example videos from different subsets of our dataset. We focus on in-the-wild audio-visual data sourced from the internet. Our dataset contains YouTube videos recorded with stereo microphones and iPhones, as well as lab-collected in-ear binaural and stereo data. We also show the estimated distribution of different sound source categories in each subset, to help visualize the contents of each dataset.
Dataset
Scene subset
Split
Size
Visibility (%)
Motion Type
IID binaural acc (%)
Clips (k)
Duration (hr)
Camera
Sound sources
StereoWalks (ours)
YT-Stereo
Train
14.4
20.0
20 / 50
Rotation&Translation
Unknown
–
Val
0 0.1
0 0.2
10 / 60
Mainly Rotation
Moving
57.5
Test
0 0.1
0 0.2
10 / 60
Mainly Rotation
Moving
57.5
Stereo-Fountain
Raw
0 5.0
0 7.0
10 / 30
Mainly Stationary
✗
76.2
Binaural-Fountain
Raw
0 1.4
0 2.0
10 / 30
Mainly Stationary
✗
98.0
Table 1 : Dataset comparison. We provide details on dataset length, the proportion of visible sound sources, camera motion types, and sound source types, with each clip representing 5 seconds. Visibility is represented by two numbers: the first indicates the number of sound sources visible for over 4 seconds within the 5-second clip, and the second denotes the number of sound sources visible for less than 4 seconds. IID Binaural acc denotes the accuracy of IID predictions respected by the actual left-right labels.
Figure 3 : Method overview. We use a camera’s egomotion to supervise binaural sound localization models. We obtain the rotation and translation of the camera using traditional methods from multi-view geometry. We then train an audio model to predict sound directions that are consistent with visual camera motions, using a dataset that includes in-the-wild walking tours.
Model
Training Set
Test Set
YT-Stereo
Sim.
YT-Stereo-iPhone
Stereo-Fountain
Simulated
MAE (°) ↓
2clf (%) ↑
8clf (%) ↑
MAE (°) ↓
2clf (%) ↑
8clf (%) ↑
MAE (°) ↓
2clf (%) ↑
8clf (%) ↑
Chance
55.3
46.7
12.7
62.1
52.0
16.0
39.4
49.0
15.0
IID – direct
–
57.5
–
–
97.0
–
–
97.4
–
GTRot [ 14 ]
✓
71.4
54.0
7.1
57.2
97.0
18.4
40.2
48.7
18.7
Ours – IID only [ 14 ]
✓
37.3
54.7
28.0
30.1
97.0
46.0
88.3
48.7
6.8
Table 2 : Comparison with state-of-the-art methods and other self-supervised methods on YT-Stereo and In-the-wild dataset. Sim. denote the simulated dataset.
Model
YT-Stereo-iPhone
Stereo-Fountain
MAE (°) ↓
2clf (%) ↑
8clf (%) ↑
MAE (°) ↓
2clf (%) ↑
8clf (%) ↑
Chance
55.3
46.7
12.7
62.1
52.0
16.0
IID-direct
–
57.5
–
–
97.0
–
GTRot [ 14 ]
71.4
54.0
7.1
55.1
97.1
18.7
Ours – Simulated
73.4
61.7
13.3
28.5
85.7
22.3
Ours – IID only [ 14 ]
37.3
54.7
28.0
29.6
97.3
46.0
Table 3: 360-Degree Sound Localization Prediction Results for YT-Stereo-iPhone and Stereo-Fountain Datasets. The YT-Stereo-iPhone dataset includes rotations and translations with mostly moving sound sources, while the Stereo-Fountain dataset has mostly stationary cameras and sound sources.
Model
(1) One Source
(2) Overlap
(3) Overlap Intermittent
(4) More Simulated Data of (3)
MAE(°) ↓
2clf(%) ↑
8clf(%) ↑
MAE(°) ↓
2clf(%) ↑
8clf(%) ↑
MAE(°) ↓
2clf(%) ↑
8clf(%) ↑
MAE(°) ↓
2clf(%) ↑
8clf(%) ↑
Chance
39.4
49.0
15.0
39.4
49.0
15.0
39.4
49.0
15.0
39.4
49.0
15.0
IID
-
97.4
-
-
80.1
-
-
80.8
-
-
80.8
-
GTRot [ 14 ]
4.3
97.5
84.7
22.5
83.5
33.5
26.9
81.5
32.5
19.7
87.2
41.0
Ours – IIDonly [ 14 ]
22.2
99.0
35.7
25.6
78.8
26.1
28.0
78.0
23.7
26.7
79.5
25.7
Ours – Full
9.8
98.3
63.4
23.2
87.2
36.2
24.1
82.1
35.7
23.7
82.4
41.2
Table 4 : Overlapping sounds make in-the-wild sound localization challenging: Experiments on three simulated datasets where the camera makes small translations in all settings. To better align with the real-world experiments, as in settings (1)(2)(3), the dataset is restricted to 30 hours, while in (4), the dataset is not limited. “Intermittent” refers to sound sources that are silent for half the time. “Overlap” refers to the presence of multiple sound sources.
Model
(1) Rotation Only
(2) Translation Only
(3) Rotation&Translation
(4) Distant Rotation
(5) Distant R&T
(6) Overlap R&T
MAE ↓
2clf ↑
8clf ↑
MAE ↓
2clf ↑
8clf ↑
MAE ↓
2clf ↑
8clf ↑
MAE ↓
2clf ↑
8clf ↑
MAE ↓
2clf ↑
8clf ↑
MAE ↓
2clf ↑
8clf ↑
Chance
39.4
49.0
15.0
39.4
49.0
15.0
39.4
49.0
15.0
39.9
49.2
15.7
39.8
49.1
15.8
39.0
49.5
15.2
IID
-
97.4
-
-
97.4
-
-
97.4
-
-
95.6
-
-
95.1
-
-
73.5
-
GTRot [ 14 ]
4.3
97.5
84.7
-
-
-
4.4
97.5
84.5
4.5
97.1
82.0
4.7
96.7
86.9
28.1
79.6
30.1
Ours – IIDonly [ 14 ]
22.2
99.0
35.7
22.7
99.0
35.7
22.0
99.1
35.9
23.1
97.8
35.2
23.0
98.1
35.0
32.5
75.0
25.7
Ours – R&B
9.8
98.3
63.4
-
-
-
10.1
98.0
60.1
10.7
98.1
59.1
11.7
98.6
58.7
25.6
81.4
34.6
Table 5 : Relationship between ego-translation and relative motion of sources. Experiments were conducted with various types of motion and distances from the camera using the simulated dataset HMSS-3D 2.0 made by SoundSpaces 2.0 [ 7 ] . We report Mean Absolute Error in the unit of degrees (°), and accuracy of 2clf and 8clf in the unit of percentage (%).
Dataset
MAE ↓
Left/Right Acc ↑
Front/Back Acc ↑
Stereo-Fountain
33.8
98.0%
51.0%
Binaural-Fountain
27.8
99.0%
69.3%
Table 6 : Evaluation of sound localization on the Stereo-Fountain and Binaural-Fountain datasets, showing improved accuracy in front/back localization with binaural recordings.
Figure 4 : Visualizations of results in the internet walking tour videos (the YT-Stereo subset). We show our predictions and ground-truth annotations with angles and audio class labels.
Figure 5 : The angle distribution of YT-Stereo-iPhone validation set. The angles range from -180 to 180 without filtering.
Dataset
Split
Size
Visibility ( % )
Motion Type
IID binaural acc ( % )
Clips
Duration
Camera
Sound sources
In the wild dataset
YT- stereo
Raw
13,000k
8.0k hrs
–
Mainly rotation
–
Train
14.6k
20hrs
-
Unknown
Unknown
YT- Stereo-iPhone
Raw
95k
80hrs
–
–
–
Train
14.6k
20hrs
20/50
Mainly rotation
Unknown
Unknown
Val
0.1k
0.2hrs
10/60
Moving
57.5
Table 7 : Dataset comparison. We provide details on dataset length, the proportion of visible sound sources, camera motion types, and sound source types, with each clip representing 5 seconds. Visibility is represented by two numbers: the first indicates the number of sound sources visible for over 4 seconds within the 5-second clip, and the second denotes the number of sound sources visible for less than 4 seconds. IID Binaural acc denotes the accuracy of IID predictions respected by the actual left-right labels.
Figure 6 : Examples from the Stereo-Music subset. We record sound using a stereo microphone in a scene containing music.
Figure 7 : The examples from the RWAVS dataset were recorded using binaural microphones in various scenarios where the sound source was a loudspeaker. Above, we present images from the original paper [ 32 ] .
Dataset
Classification Task
Accuracy (%)
RWAVS [ 32 ]
supervised 4-classification
50
Stereo-Music
supervised 4-classification
28
Table 8 : Summary of classification accuracy for RWAVS and Stereo-Music datasets.
School of Intelligence Science and Technology, Peking University · State Key Laboratory of General Artificial Intelligence, Peking University · Alibaba Token Hub, Alibaba Group +1