Conventional Japanese automatic speech recognition (ASR) is supervised by an orthographic transcript, although the same written form can correspond to different lexical readings realized in speech. Such utterances receive an identical target, so their reading distinction is absent from the supervision interface and cannot be recovered reliably by post-hoc text-only grapheme-to-phoneme conversion. We present Ruby-ASR, which refines the conventional target into a span-bound orthographic--lexical-reading sequence. Unlike separate full-sentence orthographic and phonological outputs, the ruby representation locally binds each written span to its realized reading and permits deterministic recovery of both views. We instantiate the target under subtitle-style and verbatim-style transcription conventions using a Qwen3-ASR backbone; a mora-level CTC objective provides auxiliary monotonic reading supervision. The experimental results across five Japanese benchmarks show that refining the recognition target can improve lexical-reading recovery without sacrificing readable orthographic transcription. We release the checkpoints and inference code.
Figures & tables
Interface
Orth.
Reading
Span- bound
Speech for R
Orthographic ASR
✓
–
–
N/A
ASR + G2P
✓
Inferred
–
No
Kana/phoneme ASR
–
✓
–
Yes
Separate sequence heads
✓
✓
–
Yes
Interleaved span target
✓
✓
✓
Yes
Table 1: Comparison of orthography–reading supervision interfaces. “Orth.” and “Reading” denote supervised outputs; “Span-bound” denotes explicit correspondence between each Oj and Rj ; and “Speech for R ” denotes speech-conditioned reading prediction rather than post-ASR G2P. Separate sequence targets lack explicit span correspondence, whereas Ruby-ASR uses an interleaved span-bound target.
Figure 1: Motivation and overview. Orthographic supervision assigns an identical target to acoustically distinct readings. Ruby-ASR instead binds each orthographic span to the reading realized in speech.
Model
B5K
CSJ
Book
CV8
TEDx
W.Avg.
Raw CER
Whisp.
7.20
19.77
20.28
8.27
9.52
11.36
Kotoba
8.98
18.95
22.63
9.39
10.75
12.33
Rz-NeMo
7.38
19.24
20.75
8.78
10.66
11.79
Rz-k2
6.52
16.82
20.73
7.64
9.09
10.35
Qwen
8.71
14.55
23.70
9.18
9.08
10.78
Table 2: Raw CER and script-aware CER (SA-CER, %) on five benchmarks. W.Avg. is character-count weighted. Lower is better.
Model
B5K ∗
CSJ ∗
Book
CV8
TEDx
W.Avg.
Whisp.
2.22
14.63
4.76
2.83
6.88
6.57
Kotoba
3.34
13.37
6.65
3.60
7.74
7.08
Rz-NeMo
1.89
12.65
5.85
3.18
6.88
6.17
Rz-k2
1.78
11.61
5.29
2.71
6.31
5.64
Qwen
3.07
9.90
6.85
3.50
6.14
5.74
K-Whisp.
0.87
4.39
5.39
5.27
7.31
4.66
Table 3: Kana CER (%) for lexical-reading transcription. Ruby-ASR and Kana-Whisper are scored from direct reading outputs; the remaining orthographic systems use a shared G2P. ∗ marks manually verified reading references.
Condition
Predictor input
B5K
CSJ
Oracle orth. + G2P
Ground-truth orth.
1.69
3.47
Qwen + G2P
Recognized orth.
3.07
9.90
Ruby-ASR-ver
Speech-conditioned
1.08
5.41
Ruby-ASR-sub
Speech-conditioned
1.32
5.37
Table 4: Oracle-orthography G2P diagnostic on independently verified B5K and CSJ references. Values are Kana CER (%). Oracle and Qwen hypotheses use the same G2P and normalization; Ruby-ASR uses its direct reading projection.
B5K
CSJ
Model
Kana CER
Exact
Kana CER
Exact
Oracle orth. + G2P
6.07
0.0
7.65
0.0
Qwen + G2P
6.13
17.6
12.84
5.4
Rz-k2 + G2P
4.79
22.8
14.44
3.6
Ruby-ASR-ver
1.59
57.9
8.02
16.8
Ruby-ASR-sub
2.24
50.3
8.02
17.1
Table 5: Results on G2P-hard utterances, defined by nonzero reading edit distance for oracle-orthography G2P. B5K contains 1,124/5,000 (22.5%) and CSJ 2,579/8,460 (30.5%) such utterances. Exact is utterance-level exact-reading accuracy (%).
Subset
B5K
CSJ
All correct orth.
1.71–1.94 / 97.4
1.47–1.61 / 97.2
Multi-reading
2.8–3.7 / 96.8
2.3–2.8 / ∼ 97
Rare ( 1≤f<10 )
5.3–5.5 / 86
30–32 / 23
Unseen ( f=0 )
13.0–13.2 / 70
30–31 / 24
Table 6: Conditional reading results on correctly recognized orthographic spans. Each entry is C-KanaCER / exact-reading accuracy (%). Ranges cover Ruby-ASR-ver and Ruby-ASR-sub. The all-span counts are approximately 22.5k for B5K and 19.3k for CSJ; multi-reading spans comprise about 49%.
Model
Start
End
Total
Utt.
≥ 3
Max
Whisp.
0.40
0.46
0.86
239
111
19
Kotoba
0.17
0.21
0.38
145
45
11
Rz-NeMo
0.73
0.22
0.95
296
123
19
Rz-k2
0.78
0.20
0.99
338
127
13
Ruby-ASR-sub
0.36
0.21
0.57
160
61
21
Table 7: Boundary overflow on the filtered ReazonSpeech test subset ( n=3,431 ). Start and End are normalized insertion rates (%) before and after the annotated subtitle span; Total is their sum. Utt. and ≥ 3 count utterances with any and at least three overflow characters, respectively; Max is the maximum overflow length. Lower is better.
Automatic speech recognition (ASR) has achieved substantial gains in transcription accuracy, yet verbatim transcription does not necessarily produce readily usable text. It retains fillers, repetitions, false starts, and self-corrections that increase reading effort, obscure the speaker's final intent, and propagate unresolved or abandoned content to downstream tasks. Existing spoken-to-written methods process completed audio or transcripts but cannot revise emitted text when later speech changes how preceding content should be interpreted. We therefore formulate Agentic Speech Recognition (AgenticSR), an audio-to-clean-text task that removes disfluencies, resolves self-corrections, and normalizes written form while preserving the speaker's final intent. AgenticASR implements this task through an ASR--Refiner architecture that repeatedly transforms a bounded active context and replaces its corresponding output span as audio arrives. This enables continual emission and revision over streams of arbitrary duration. We also introduce AASR-Bench, a bilingual benchmark with fine-grained atomic rubrics. Across multiple ASR front ends, AgenticASR attains the highest AASR-Bench scores among evaluated systems. A human--AI agreement study shows that rubric-based judgments align with independent expert assessments. Ablations characterize Refiner capacity, context length, and the quality--latency trade-off between online and offline inference. Together, these results establish AgenticASR as a practical framework for intent-preserving clean transcription during ongoing speech. Code, AASR-Bench, and a demo will be released at https://github.com/AnXMuy/AgenticASR.
Zixuan Jiang, Binghao Qiang, Jiaying Chi +3
X-LANCE lab, Shanghai Jiao Tong University · Shanghai Innovation Institute · College of Artificial Intelligence, Xi’an Jiaotong University
Automatic speech recognition systems commonly rely on reference transcriptions for evaluation, while reference-free approaches often depend on internal confidence estimation or auxiliary language models. We propose READ (Reference-free Hypothesis Evaluation with Acoustic Discrepancy), a novel metric that evaluates ASR hypotheses directly from the speech signal. READ emphasizes the acoustic grounding of hypotheses. It uses a pretrained auto-regressive TTS model to compute the conditional likelihood of speech tokens given a text hypothesis, to measure fine-grained acoustic discrepancy between speech and text. Without additional training, READ can be applied for hypothesis refinement. Experiments show that READ correlates with specific recognition errors and improves ASR outputs, achieving up to 20% relative error rate reduction, with particularly strong gains under noisy conditions.
Zhihan Li, Hankun Wang, Yiwei Guo +3
X-LANCE Lab, School of Computer Science, Shanghai Jiao Tong University, China
ASR-roundtrip evaluation is widely used as a scalable proxy for text-to-speech (TTS) intelligibility, but it can produce false negatives for reading errors perceived by listeners. We study Chinese news TTS spans whose correct reading depends on context or domain conventions, such as sports scores, aircraft models, technical units, and membership names. In these cases, Raw TTS can choose a plausible but wrong reading while ASR transcribes the audio as the intended or surface-correct text. A targeted audit over 110 high-risk MiMo TTS cases, reported with a complete denominator, confirms 46 masked false negatives, 9 exposed TTS errors, and 55 cases with no Raw TTS error. A span-isolation diagnostic re-exposes 18/46 previously masked errors. A Raw-only CosyVoice audit on the same targeted pool confirms 51 masked cases. Across the 97 TTS-specific audio files labeled confirmed masked across the two audits, Qwen3-ASR surface-recovers 40 cases, whereas Paraformer does so in only 2. The results suggest that ASR-roundtrip is useful for screening but insufficient as standalone ground truth for Chinese news reading-risk evaluation.