Far-field automatic speech recognition(ASR) degrades under reverberation, noise, and talker motion, yet the benchmarks that drive model selection emphasize close-microphone speech. We present FFASR, a held-out corpus of 15,637 utterances and an open leaderboard spanning nine conditions, each varying a single acoustic factor: anechoic near-field speech, a measured-versus-simulated office-lab pair, static far-field mixtures at high/mid/low signal-to-noise ratio(SNR), and moving-talker variants at matched SNR. Dry speech from 15 talkers is convolved with hybrid wave/geometrical-acoustics room impulse responses from 14 furnished rooms; because the speech is newly recorded and the test waveforms are never released, the corpus resists training-data contamination. Across contemporary systems, mean word error rate (WER) rises from 4.4% near-field to 41.3% in the static low-SNR condition; a moving talker adds a small but consistent penalty at matched SNR; and on the office-lab pair, measured and simulated WER agree to within about 1.7 pp on average. These results support high-fidelity simulation as a scalable proxy for measured far-field evaluation under the conditions we test.
Figures & tables
Fig. 1: Data-generation pipeline. Newly recorded anechoic speech is convolved with hybrid wave/GA room impulse responses and seeded noise draws; each mixture is labelled by its post-render SNR.
Condition
N
Description
Near field
889
Dry anechoic speech
Lab simulated
2000
Stage A hybrid render
Lab measured
2000
Stage A measured mono
High SNR
1746
Static, SNR ≥14 dB
Mid SNR
1938
Static, 8–12 dB
Low SNR
1647
Static, ≤6 dB
TABLE I: FFASR evaluation conditions. N is the number of clips in each packed split.
Quantity
Min
Max
Mean
SD
Room volume (m 3 )
17.47
367.41
128.52
77.26
T20 (s)
0.19
1.29
0.60
0.3
Source–receiver dist. (m)
0.3
13.8
3.98
2.2
C50 (dB)
-7.8
19.61
5.3
4.5
Post-render SNR (dB)
-1.9
88.9
10.5
6.4
TABLE II: Acoustic statistics of the FFASR furnished-room renders.
Fig. 2: Example scene from the dataset. Left: the target speech source, the directional transient and ambient noise sources, and the receiver. Right: an example trajectory of the moving speech source.
Condition
Mean
Med
SD
Min
Max
Near field
4.4
4.3
0.4
3.8
5.4
Lab measured
29.0
27.8
6.7
20.0
45.0
Lab simulated
27.7
26.9
6.9
18.6
42.6
Static High SNR
10.1
10.2
2.2
6.7
15.1
Static Mid SNR
20.5
20.3
4.7
13.9
29.5
Static Low SNR
41.3
40.2
8.3
28.4
57.0
TABLE III: Distribution of WER (%) across the 15 leading evaluated systems, summarized per condition. The lower panel reports the two derived effects in percentage points (pp).
Fig. 3: Average WER vs. RTFx on NVIDIA L4 (batch size 1). Efficient CTC/TDT models occupy the upper-left region.