Organizations: ETH Zurich · Sony Corporate Technology Center America, Inc. · Sony Europe Ltd. · Acoustics Lab, DICE, Aalto University · Sony Group Corporation
Historical music restoration (HMR) has almost exclusively focused on constrained problems such as Super-Resolution or the restoration of solo pieces, under-exploring the general task of restoring orchestral historical music, which has multiple instruments. This under-exploration is largely because the HMR domain, early-20th-century recordings, has no pre-degradation ground-truth pairs, making the restoration task unsupervised and more challenging. This paper presents a supervised end-to-end orchestral HMR benchmark by exploring both the synthetic degradation functions and the end-to-end generative deep-learning restoration methods. We simulate the historical recording degradation chain more faithfully than prior work, which makes orchestral restoration into a tractable supervised problem. A latent flow-matching model trained on the resulting synthetic pairs outperforms existing HMR baselines on intrusive, non-intrusive, and subjective evaluations. We also curate and release a 9.3-hour license-free, unpaired, historical classical-music test set, along with code and audio demos.
Figures & tables
Figure 1 : Overview of the proposed historical music restoration pipeline Rθ,CFM .
MERT [ 10 ]
Fx-Enc. [ 33 ]
SAME-L [ 27 ]
Degradation
Full Orch.
Light Orch.
Full Orch.
Light Orch.
Full Orch.
Light Orch.
Clean
5.39
5.40
5.43
5.49
5.33
5.45
Gaussian Noise
5.18
5.13
4.71
5.08
5.40
5.44
Low-Pass + Gaussian Noise
5.31
5.46
5.42
5.49
5.24
5.37
Gramophone Noise
3.90
4.42
3.86
4.63
4.57
4.82
WH + Gramophone Noise (Ours)
4.23
4.55
4.42
4.99
4.54
4.95
Table 1 : MAD ( ↓ ) between degradation ablations embeddings on FOS and historical unpaired test set. The final row directly compares the two subsets of the unpaired test set.
Codec
Paired Clean
Paired 5-Stage
Full Orchestra
Light Orchestra
CoDiCodec [ 28 ]
1.08±0.45
0.25±0.27
1.37±0.68
1.24±0.69
SAO [ 7 ]
0.39±0.19
0.12±0.17
0.80±0.50
0.88±0.58
DAC [ 14 ]
0.22±0.10
0.08±0.15
0.72±0.57
0.87±0.63
SAME-L
0.11±0.08
0.05±0.12
0.62±0.57
0.81±0.62
Table 2 : Autoencoder reconstruction Mel-MSE with equal-window aggregation. Waveforms were encoded and decoded by the corresponding autoencoder and compared against the input using Mel-MSE
Method
Params
Historical Unpaired Test Set
Synthetic Paired Test Set
AA-PQ ( ↑ )
MOS-P ( ↑ )
MOS-Q ( ↑ )
AA-PQ ( ↑ )
Mel MSE ( ↓ )
CLAP Cos. ( ↑ )
MAD ( ↓ )
Full-Orch.
Light-Orch.
Unprocessed input
–
4.94
5.00
–
1.17
5.74
3.94
0.44
5.03
Ground truth
–
–
–
–
4.44
7.44
0.00
1.00
0.00
BEHM-GAN
82M
4.87
4.93
–
–
6.05
5.33
0.45
4.94
BABE2
40M
5.88
5.87
3.70
3.52
6.35
4.99
0.56
5.46
Table 3 : Restoration performance on objective and subjective evaluations.
Figure 2 : MOS-Q and MOS-P means, 95% confidence intervals, and non-outlier ranges
Music restoration seeks to recover a clean signal from an observed recording degraded by an unknown effect, distortion, or corruption. Existing systems often rely on paired training data and distortion-specific supervision, which limits their use when the forward process is not known in advance. We propose LOUDAR (Latent-space Optimization of Unknown Distortion for Audio Restoration) a general-purpose restoration method that operates in the latent space of a pretrained audio autoencoder and models the unknown distortion as a learnable latent operator. At inference time, LOUDAR alternates between estimating the clean latent variable and updating the latent operator parameters. An unconditional latent diffusion model provides a prior over clean audio and regularizes this inference by steering the latent estimate toward the manifold of clean recordings. Because the degradation model is adapted per input, the approach is broadly applicable across diverse restoration problems. We evaluate LOUDAR on singing voice effect removal and restoration, as well as guitar distortion removal, and show that it consistently improves over degraded inputs and is competitive with supervised and unsupervised baselines in waveform and latent domains.
Michal Švento, Eloi Moliner, Valtteri Kallinen +3
Dept. of Telecommunications, Brno University of Technology, Czech Republic · Acoustic Lab, Dept. of Information and Communications Engineering, Aalto University, Finland
Recent audio restoration increasingly relies on large-scale conditional latent generative modeling, including diffusion, Schrodinger Bridges, and Flow Matching variants, to invert degradations such as bandwidth limitation or noise. We present an analysis of the performance of various state-of-the-art methods compared to simple arithmetic transformations in the latent spaces of multiple neural codecs for musical bandwidth extension. We show that estimating a single transport vector between the clean and degraded latent centroids on a reference set, and adding it to degraded latents, can yield restoration performance competitive with large diffusion models. This suggests, first, that some neural codec latent spaces exhibit structure aligned with audio bandwidth; and second, that in such cases complex conditional models may offer only limited gains over a simple vector addition. We argue that these findings reveal an interesting avenue for future research whereby models could take advantage of the latent space structure in order to offer greater training and parameter efficiency, and overall better performance. Additionally, we propose to consider this simple arithmetic transformation as a baseline for music bandwidth extension research, as it allows an assessment of the contribution of learnable parameters towards restoration performance.
Hendrik Vincent Koops, Hao Hao Tan, Elio Quinton
Music & Audio Machine Learning Lab · Universal Music Group, London, U.K.
Archival footage often suffers from coupled visual and acoustic degradations, yet most restoration systems process the two modalities separately. To address this problem, we present OmniVR, the first systematic framework for joint audio-video restoration, covering data construction, model adaptation, efficient inference, and evaluation. We construct a high-quality audio-video corpus with detailed captions and use a joint degradation pipeline to produce aligned clean and degraded pairs. Using these pairs, we adapt a pretrained text-to-audio-video model (T2AV) by introducing degraded audio-video conditions (TAV2AV), then progressively replace sample captions with a fixed restoration prompt while retaining caption/null rehearsal. The resulting AV2AV model requires no user-provided text. Under a compatible residual-learning model, we prove that this condition-annealing schedule reduces gradient variance and expected restoration risk relative to direct fixed-prompt adaptation at the same training budget. For efficient deployment, OmniVR-Flash combines reduced-resolution video conditioning, MeanFlow-based one-step distillation, and Turbo VAE, achieving approximately 38 fps at 1K and 18 fps at 2K on a single B200 GPU. We further introduce OmniVRBench to evaluate four complementary dimensions: visual quality, audio quality, temporal consistency, and audio-visual synchrony. OmniVR achieves state-of-the-art results on public benchmarks and OmniVRBench. Data, code, and model weights will be released. Project Page: https://xin1u.github.io/OminiVR_PAGE/
Xin Lu, Zihao Fan, Jie Huang +4
University of Science and Technology of China · JD Explore Academy