Supervising Sound Localization by In-the-wild Egomotion
Organizations: Tsinghua University · University of Michigan · Shanghai Qi Zhi Institute
Abstract
We present a method for learning binaural sound localization using egomotion as a supervisory signal. Over the course of a video, the cameras direction to a sound source will change as the camera moves. We train an audio model to predict sound directions that are consistent with visual estimates of camera motion, which we obtain using traditional methods from multi-view geometry. This provides a weak but plentiful form of supervision that we combine with traditional binaural cues. To evaluate this method, we propose a dataset of real-world audio-visual videos with egomotion. We show that our model can successfully learn from real-world data and that it performs well on sound localization tasks
Figures & tables
| Dataset | Scene subset | Split | Size | Visibility (%) | Motion Type | IID binaural acc (%) | ||
| Clips (k) | Duration (hr) | Camera | Sound sources | |||||
| StereoWalks (ours) | YT-Stereo | Train | 14.4 | 20.0 | 20 / 50 | Rotation&Translation | Unknown | – |
| Val | 0 0.1 | 0 0.2 | 10 / 60 | Mainly Rotation | Moving | 57.5 | ||
| Test | 0 0.1 | 0 0.2 | 10 / 60 | Mainly Rotation | Moving | 57.5 | ||
| Stereo-Fountain | Raw | 0 5.0 | 0 7.0 | 10 / 30 | Mainly Stationary | ✗ | 76.2 | |
| Binaural-Fountain | Raw | 0 1.4 | 0 2.0 | 10 / 30 | Mainly Stationary | ✗ | 98.0 | |
| Model | Training Set | Test Set | |||||||||
| YT-Stereo | Sim. | YT-Stereo-iPhone | Stereo-Fountain | Simulated | |||||||
| MAE (°) | 2clf (%) | 8clf (%) | MAE (°) | 2clf (%) | 8clf (%) | MAE (°) | 2clf (%) | 8clf (%) | |||
| Chance | 55.3 | 46.7 | 12.7 | 62.1 | 52.0 | 16.0 | 39.4 | 49.0 | 15.0 | ||
| IID – direct | – | 57.5 | – | – | 97.0 | – | – | 97.4 | – | ||
| GTRot [ 14 ] | ✓ | 71.4 | 54.0 | 7.1 | 57.2 | 97.0 | 18.4 | 40.2 | 48.7 | 18.7 | |
| Ours – IID only [ 14 ] | ✓ | 37.3 | 54.7 | 28.0 | 30.1 | 97.0 | 46.0 | 88.3 | 48.7 | 6.8 | |
| Model | YT-Stereo-iPhone | Stereo-Fountain | |||||
| MAE (°) | 2clf (%) | 8clf (%) | MAE (°) | 2clf (%) | 8clf (%) | ||
| Chance | 55.3 | 46.7 | 12.7 | 62.1 | 52.0 | 16.0 | |
| IID-direct | – | 57.5 | – | – | 97.0 | – | |
| GTRot [ 14 ] | 71.4 | 54.0 | 7.1 | 55.1 | 97.1 | 18.7 | |
| Ours – Simulated | 73.4 | 61.7 | 13.3 | 28.5 | 85.7 | 22.3 | |
| Ours – IID only [ 14 ] | 37.3 | 54.7 | 28.0 | 29.6 | 97.3 | 46.0 | |
| Model | (1) One Source | (2) Overlap | (3) Overlap Intermittent | (4) More Simulated Data of (3) | ||||||||
| MAE(°) | 2clf(%) | 8clf(%) | MAE(°) | 2clf(%) | 8clf(%) | MAE(°) | 2clf(%) | 8clf(%) | MAE(°) | 2clf(%) | 8clf(%) | |
| Chance | 39.4 | 49.0 | 15.0 | 39.4 | 49.0 | 15.0 | 39.4 | 49.0 | 15.0 | 39.4 | 49.0 | 15.0 |
| IID | - | 97.4 | - | - | 80.1 | - | - | 80.8 | - | - | 80.8 | - |
| GTRot [ 14 ] | 4.3 | 97.5 | 84.7 | 22.5 | 83.5 | 33.5 | 26.9 | 81.5 | 32.5 | 19.7 | 87.2 | 41.0 |
| Ours – IIDonly [ 14 ] | 22.2 | 99.0 | 35.7 | 25.6 | 78.8 | 26.1 | 28.0 | 78.0 | 23.7 | 26.7 | 79.5 | 25.7 |
| Ours – Full | 9.8 | 98.3 | 63.4 | 23.2 | 87.2 | 36.2 | 24.1 | 82.1 | 35.7 | 23.7 | 82.4 | 41.2 |
| Model | (1) Rotation Only | (2) Translation Only | (3) Rotation&Translation | (4) Distant Rotation | (5) Distant R&T | (6) Overlap R&T | ||||||||||||
| MAE | 2clf | 8clf | MAE | 2clf | 8clf | MAE | 2clf | 8clf | MAE | 2clf | 8clf | MAE | 2clf | 8clf | MAE | 2clf | 8clf | |
| Chance | 39.4 | 49.0 | 15.0 | 39.4 | 49.0 | 15.0 | 39.4 | 49.0 | 15.0 | 39.9 | 49.2 | 15.7 | 39.8 | 49.1 | 15.8 | 39.0 | 49.5 | 15.2 |
| IID | - | 97.4 | - | - | 97.4 | - | - | 97.4 | - | - | 95.6 | - | - | 95.1 | - | - | 73.5 | - |
| GTRot [ 14 ] | 4.3 | 97.5 | 84.7 | - | - | - | 4.4 | 97.5 | 84.5 | 4.5 | 97.1 | 82.0 | 4.7 | 96.7 | 86.9 | 28.1 | 79.6 | 30.1 |
| Ours – IIDonly [ 14 ] | 22.2 | 99.0 | 35.7 | 22.7 | 99.0 | 35.7 | 22.0 | 99.1 | 35.9 | 23.1 | 97.8 | 35.2 | 23.0 | 98.1 | 35.0 | 32.5 | 75.0 | 25.7 |
| Ours – R&B | 9.8 | 98.3 | 63.4 | - | - | - | 10.1 | 98.0 | 60.1 | 10.7 | 98.1 | 59.1 | 11.7 | 98.6 | 58.7 | 25.6 | 81.4 | 34.6 |
| Dataset | MAE | Left/Right Acc | Front/Back Acc |
| Stereo-Fountain | 33.8 | 98.0% | 51.0% |
| Binaural-Fountain | 27.8 | 99.0% | 69.3% |
| Dataset | Split | Size | Visibility ( % ) | Motion Type | IID binaural acc ( % ) | |||
| Clips | Duration | Camera | Sound sources | |||||
| In the wild dataset | YT- stereo | Raw | 13,000k | 8.0k hrs | – | Mainly rotation | – | |
| Train | 14.6k | 20hrs | - | Unknown | Unknown | |||
| YT- Stereo-iPhone | Raw | 95k | 80hrs | – | – | – | ||
| Train | 14.6k | 20hrs | 20/50 | Mainly rotation | Unknown | Unknown | ||
| Val | 0.1k | 0.2hrs | 10/60 | Moving | 57.5 | |||
| Dataset | Classification Task | Accuracy (%) |
| RWAVS [ 32 ] | supervised 4-classification | 50 |
| Stereo-Music | supervised 4-classification | 28 |