eess.ASAug 1, 2026

Simulation-Based Plate-Reverb Parameter Estimation from a Single Impulse Response

Authors: Minhui LuJoshua D. Reiss

Organizations: Centre for Digital Music Queen Mary University of London London, United Kingdom

Abstract

We present a simulation-trained, non-iterative estimator for Task A of the 1st DAFx Parameter Estimation Challenge. Each unnormalized plate-reverb impulse response is summarized by amplitude, spectral, and decay descriptors, and an ensemble of tree regressors estimates the six target parameters in one pass. Across two independent synthetic validation sets, the normalized models outperform the training-set mean and an earlier raw-regression baseline. On a shared set, the final ensemble also outperforms a single run of the official default PSO at substantially lower inference cost. Since the official labels are hidden, parameter accuracy is measured on simulator-matched data, and the released responses support only audio-side consistency checks. The estimator returns point estimates without uncertainty.

Explore similar work

Aug 1, 2026eess.AS

Band-Count Dense Modal Estimation with Fixed-Frequency Differentiable Resonator Refinement

Task B of the 1st DAFx Parameter Estimation Challenge requires estimating the frequencies, decay rates, gains, and number of modes in a dense plate-reverb impulse response. Weak and overlapping modes make sparse peak detection prone to severe undercounting. We train an ExtraTrees regressor on simulator-generated data to predict mode counts in four frequency bands. These counts define dense frequency grids, after which a differentiable all-pole resonator model refines decay and gain while keeping frequency fixed. On two separate synthetic validation sets, the system reduces a local challenge-style error by about 66% relative to the official default peak-picking baseline. The improvement is mainly associated with lower mode-count mismatch, while decay and gain remain the largest error sources. These findings support separating modal-density estimation from continuous parameter fitting.
Minhui Lu, Joshua D. Reiss
May 8, 2026eess.AS

Dependence on Early and Late Reverberation of Single-Channel Speaker Distance Estimation

Single-channel speaker distance estimation has recently achieved centimeter-level accuracy in simulated environments, yet it remains unclear which components of the room impulse response (RIR) the model exploits and how performance depends on the recording conditions. In this work, we decompose simulated RIRs into four variants (full, direct-only, no-late, and no-early) using the mixing time estimated from the echo density function as the boundary between early reflections and late reverberation. We define four calibration scenarios, from fully calibrated (synchronised capture, known source level) to fully uncalibrated (arbitrary onset, unknown level), and evaluate all combinations on a matched dataset. Results show that without time calibration, mean absolute error (MAE) increases to 1.291.29 m and the model extracts reverberation-based cues, with early reflections emerging as the most informative component. Further analysis against DRR, C50C_{50}, and T60T_{60} confirms that estimation accuracy improves with stronger early energy and degrades in highly reverberant environments. When time calibration is available, the model achieves a MAE of 0.140.14 m by extracting the propagation delay alone, regardless of the RIR content.
Michael Neri, Archontis Politis, Tuomas Virtanen
Date pendingeess.AS

What does the model actually see? Evaluation protocols and input availability in data-driven prediction of room acoustic parameters

Machine-learnt models are increasingly used to predict ISO 3382-1 room acoustic parameters at unmeasured seats from sparse measurements, with reported coefficients of determination frequently above 0.85. This paper shows that such figures are often determined by the evaluation protocol rather than by the model. Using a multi-condition measurement campaign in a 264-seat conference hall and a 180-seat concert hall, three model families were evaluated under a factorial protocol ablation: validation splits either row-based or grouped by receiver position, and inputs either including measured-at-test quantities (the target position's impulse response, co-located parameter measurements, position identifiers) or restricted to source-receiver geometry and environmental state. Row-based splits with measured-at-test inputs reproduce the high reported accuracies (mean R^2 of 0.81 for the core parameters); grouped splits with deployment-consistent inputs reduce these to 0.09-0.60 and reorder the apparent difficulty of parameter classes. Access to the target's own impulse response at test time reproduces the high accuracy of the row-based folds but provides no benefit under grouped folds. The row-based figure reflects condition interpolation at the measured positions rather than transferable acoustic information. Under the deployment-consistent protocol the spread between Random Forest, the hybrid network, and inverse-distance weighting is several times smaller than the spread between protocols for a fixed model; the learnt models retain a measurable advantage for sound strength and reverberation time, and the high accuracy of the original pipelines re-emerges as condition interpolation at measured positions, a distinct and operationally useful task. A reporting checklist operationalises the findings.
Ak\in Oktav