Far-field automatic speech recognition(ASR) degrades under reverberation, noise, and talker motion, yet the benchmarks that drive model selection emphasize close-microphone speech. We present FFASR, a held-out corpus of 15,637 utterances and an open leaderboard spanning nine conditions, each varying a single acoustic factor: anechoic near-field speech, a measured-versus-simulated office-lab pair, static far-field mixtures at high/mid/low signal-to-noise ratio(SNR), and moving-talker variants at matched SNR. Dry speech from 15 talkers is convolved with hybrid wave/geometrical-acoustics room impulse responses from 14 furnished rooms; because the speech is newly recorded and the test waveforms are never released, the corpus resists training-data contamination. Across contemporary systems, mean word error rate (WER) rises from 4.4% near-field to 41.3% in the static low-SNR condition; a moving talker adds a small but consistent penalty at matched SNR; and on the office-lab pair, measured and simulated WER agree to within about 1.7 pp on average. These results support high-fidelity simulation as a scalable proxy for measured far-field evaluation under the conditions we test.
Figures & tables
Fig. 1: Data-generation pipeline. Newly recorded anechoic speech is convolved with hybrid wave/GA room impulse responses and seeded noise draws; each mixture is labelled by its post-render SNR.
Condition
N
Description
Near field
889
Dry anechoic speech
Lab simulated
2000
Stage A hybrid render
Lab measured
2000
Stage A measured mono
High SNR
1746
Static, SNR ≥14 dB
Mid SNR
1938
Static, 8–12 dB
Low SNR
1647
Static, ≤6 dB
TABLE I: FFASR evaluation conditions. N is the number of clips in each packed split.
Quantity
Min
Max
Mean
SD
Room volume (m 3 )
17.47
367.41
128.52
77.26
T20 (s)
0.19
1.29
0.60
0.3
Source–receiver dist. (m)
0.3
13.8
3.98
2.2
C50 (dB)
-7.8
19.61
5.3
4.5
Post-render SNR (dB)
-1.9
88.9
10.5
6.4
TABLE II: Acoustic statistics of the FFASR furnished-room renders.
Fig. 2: Example scene from the dataset. Left: the target speech source, the directional transient and ambient noise sources, and the receiver. Right: an example trajectory of the moving speech source.
Condition
Mean
Med
SD
Min
Max
Near field
4.4
4.3
0.4
3.8
5.4
Lab measured
29.0
27.8
6.7
20.0
45.0
Lab simulated
27.7
26.9
6.9
18.6
42.6
Static High SNR
10.1
10.2
2.2
6.7
15.1
Static Mid SNR
20.5
20.3
4.7
13.9
29.5
Static Low SNR
41.3
40.2
8.3
28.4
57.0
TABLE III: Distribution of WER (%) across the 15 leading evaluated systems, summarized per condition. The lower panel reports the two derived effects in percentage points (pp).
Fig. 3: Average WER vs. RTFx on NVIDIA L4 (batch size 1). Efficient CTC/TDT models occupy the upper-left region.
Despite rapid advances in automatic speech recognition (ASR) and large audio-language models, robust recognition in real-world environments remains limited by an "acoustic robustness bottleneck": models often lose acoustic grounding and produce omissions or hallucinations under severe, compositional distortions. We propose Mega-ASR, a unified ASR-in-the-wild framework that combines scalable compound-data construction with progressive acoustic-to-semantic optimization. We introduce Voices-in-the-Wild-2M, covering 7 classic acoustic phenomena and 54 physically plausible compound scenarios, and train Mega-ASR with Acoustic-to-Semantic Progressive Supervised Fine-Tuning and Dual-Granularity WER-Gated Policy Optimization. Extensive experiments demonstrate that Mega-ASR achieves significant advantages over prior state-of-the-art systems on adverse-condition ASR benchmarks (45.69% vs. 54.01% on VOiCES R4-B-F, and 21.49% vs. 29.34% on NOIZEUS Sta-0). On complex compositional acoustic scenarios, Mega-ASR further delivers over 30% relative WER reduction against strong open- and closed-source baselines, establishing a scalable paradigm for robust ASR in-the-wild.
While automatic speech recognition (ASR) models have achieved remarkable improvements in recent years, performance disparities persist across different speaker populations. One such disparity is for speakers whose first languages (L1) are from families distant from English. This paper investigates the relationship between first language background and English ASR performance. Through empirical analysis, we observe that the correlation between speakers' L1 distance and ASR error rates yields a systematic effect on English Speech, with its strength varying across datasets and models. This association is statistically significant in a follow-up analysis accounting for dataset-level variation in Tweedie mixed-effects models (p<0.001 across evaluated models). In addition, analysis of the latent space reveals a L1-based spatial segregation across deeper acoustic layers in the majority of evaluated architectures
Ting-Hui Cheng, Line Katrine Harder Clemmensen, Sneha Das
Department of Applied Mathematics and Computer Science Technical University of Denmark
Room-acoustic simulations are widely used to augment training data for deep-learning-based speech enhancement. While most pipelines rely on simplified geometrical acoustics, wave-based approaches offer greater physical accuracy. In this work, we examine how simulation fidelity affects multichannel speech enhancement performance. To this end, we train SpatialNet on datasets augmented with different room-acoustic simulation methods and evaluate the resulting models on measured data. We compare lower-fidelity datasets based on geometrical acoustics with a high-fidelity dataset using advanced acoustic modelling and a hybrid combination of wave-based and geometrical acoustics simulations. Training on the high-fidelity dataset results in an up to 38 % relative reduction in median word error rate compared to the lower-fidelity alternatives. These results show that augmentation with high-fidelity room-acoustic simulations directly translates into improved multichannel speech enhancement performance.