eess.ASSep 24, 2026

Few-Shot Calibration for Sim-to-Real Single-Channel Speaker Distance Estimation

Authors: Michael Neri, Archontis Politis, Tuomas Virtanen

Organizations: Faculty of Information Technology and Communication Sciences, Tampere University, Finland

Abstract

Speaker distance estimators are trained almost exclusively on simulated room acoustics, because real recordings annotated with the true talker-to-microphone distance are scarce. We show that models trained this way transfer poorly. On three real corpora we evaluate, simply predicting the average distance of the corpus is more accurate than any learned model. Then, we ask how few labelled real utterances are needed to make a frozen, synthetic-trained estimator useful, and study post-hoc calibration maps that rescale its output without gradients or retraining. An analysis of the achievable error shows that what the calibration is not limited by the absolute accuracy of the estimator, but how well it orders utterances by distance, since a constant bias or a wrong output scale is removed exactly by the calibration itself. Balancing this against the cost of estimating each coefficient from few samples yields a criterion that accounts for which map wins on which corpus and at which annotation budget, together with a shrinkage variant that requires no hard decision. Our findings suggest selecting synthetic checkpoints by linear correlation with true distances rather than by absolute error. Code, datasets, and analysis are available at https://github.com/michaelneri/audio-distance-estimation.

Figures & tables

Explore similar work

May 8, 2026eess.AS

Dependence on Early and Late Reverberation of Single-Channel Speaker Distance Estimation

Single-channel speaker distance estimation has recently achieved centimeter-level accuracy in simulated environments, yet it remains unclear which components of the room impulse response (RIR) the model exploits and how performance depends on the recording conditions. In this work, we decompose simulated RIRs into four variants (full, direct-only, no-late, and no-early) using the mixing time estimated from the echo density function as the boundary between early reflections and late reverberation. We define four calibration scenarios, from fully calibrated (synchronised capture, known source level) to fully uncalibrated (arbitrary onset, unknown level), and evaluate all combinations on a matched dataset. Results show that without time calibration, mean absolute error (MAE) increases to 1.291.29 m and the model extracts reverberation-based cues, with early reflections emerging as the most informative component. Further analysis against DRR, C50C_{50}, and T60T_{60} confirms that estimation accuracy improves with stronger early energy and degrades in highly reverberant environments. When time calibration is available, the model achieves a MAE of 0.140.14 m by extracting the propagation delay alone, regardless of the RIR content.
May 1, 2026cs.SD

Towards Improving Speaker Distance Estimation through Generative Impulse Response Augmentation

The Room Acoustics and Speaker Distance Estimation (SDE) Challenge at ICASSP 2025 explores the effectiveness of augmented room impulse response (RIR) data for improving SDE model performance. This challenge at GenDARA involves generating RIRs to supplement sparse datasets and fine-tuning SDE models with the augmented data. We employ the open-source fast diffuse room impulse response generator (FastRIR) conditioned only on speaker and listener locations. We design a quality filter to ensure generated RIR alignment with challenge RIRs, and hyperparameter optimization is employed for model fine-tuning. Our approach reduces the mean absolute error (MAE) of the five positions from 1.66m to 0.6m for GWA rooms and from 2.18m to 0.69m for Treble rooms, with results demonstrating that the augmentation approach significantly improves estimation accuracy, particularly at medium to long distances.
Jun 10, 2026cs.SD

Fast-SDE: Efficient Single-Microphone Sound Source Distance Estimation in Reverberant Environments

Sound source distance estimation (SDE) is a critical capability in human-robot interaction. An inappropriate interaction distance not only reduces the reliability of speech acquisition and understanding, but also compromises the naturalness and comfort of the interaction. Most existing SDE methods rely on microphone arrays, however, multi-microphone systems typically require careful hardware synchronization, geometric calibration, and additional space and computational resources, which limits applicability to size-constrained and computability-limited embodied platforms. To alleviate these issues, we propose Fast-SDE, a lightweight single-microphone SDE framework that is suited for deployment on robot platforms with limited computational resources and strict size constraints. Specifically, Fast-SDE employs a subband-based backbone that decomposes the frequency axis into multiple subbands, rather than processing the entire spectrum with a wide full-band backbone. A shared subband encoder then maps each subband to a compact latent representation and learns the relationship between acoustic structure and time-frequency patterns. Finally, a lightweight regression head converts the fused subband representations into the estimated distance. Extensive simulation and real-world experiments demonstrate the merits of the proposed method. To benefit the broader research community, we have open-sourced our code at https://github.com/JiangWAV/FAST-SDE.