Recent advances in speech language models have improved automatic speech recognition (ASR) for long-form audio. However, accurately and consistently transcribing domain-specific terminology remains challenging. Motivated by the world knowledge and contextual capability of large language models (LLMs), we propose Agentic-GER, an LLM-based agent for terminology correction in long-form speech. The agent uses global context from the full transcript to identify suspicious terms and resolve ambiguous hypotheses. It selectively re-transcribes the source speech to check candidate corrections, and uses accepted edits to guide subsequent decisions. Experiments with four LLMs and two ASR systems on GigaSpeechBench show consistent terminology improvements in both Chinese and English, with and without thinking. On Chinese speech, Agentic-GER achieves up to a 36.8% relative reduction in biased character error rate (B-CER) over the Whisper baseline.
Figures & tables
Figure 1: Overview of Agentic-GER. Global context guides error discovery and correction. The agent re-transcribes selected segments to evaluate edits, updates the transcript, and records decisions for subsequent rounds.
Figure 2: Resolving an ambiguous ASR hypothesis using global context. The agent flags “death certificates” in a discussion of medical device evolution. Re-transcription yields “deathscapes.” An earlier segment mentions a “stethoscope.” The agent corrects the term to “stethoscopes,” matching the reference. Excerpts are shortened for display.
Chinese
English
B-CER ↓
B-WER ↓
FunASR
Whisper
FunASR
Whisper
Baseline
10.91
35.07
12.94
14.67
+Qwen3.8-27B
10.19 (6.7)
28.88 (17.7)
12.35 (4.6)
14.08 (4.0)
+Gemma4-31B
9.88 (9.5)
25.09 (28.5)
12.05 (6.9)
13.71 (6.5)
+Qwen3.8-Flash
9.64 (11.7)
26.61 (24.1)
11.91 (8.0)
13.72 (6.5)
Table 1: Main terminology results without thinking. Error rates (%, ↓ ); parentheses show relative error reductions (%, ↑ ) from Baseline. Column minima are in bold.
Language
Error type
Proportion
Reduction ↑
Chinese
Substitution
85.0
21.3
Deletion
14.7
3.0
English
Substitution
56.0
6.4
Deletion
43.1
0.4
Table 2: Terminology error analysis for Qwen3.8-27B on Whisper without thinking. Proportion (%) is each error type’s percentage of all baseline terminology errors, including insertions. Reduction (%, ↑ ) is the relative decrease in its error count.
Thinking
B-CER ↓
CER ↓
Baseline
—
10.91
3.12
yq +Local GER
✗
10.80
3.37
✓
10.63
3.42
yq +Agentic-GER
✗
10.19
3.11
✓
9.03
3.01
Table 3: Local GER versus Agentic-GER. Error rates (%, ↓ ); best values in bold. “Thinking” indicates whether intermediate token generation is enabled ( ✓ ) or disabled ( ✗ ).
Chinese
English
B-CER ↓
B-WER ↓
FunASR
Whisper
FunASR
Whisper
+Qwen3.8-27B
9.03
25.89
11.96
13.87
Δ Thinking (pp) ↑
+10.6
+8.5
+3.0
+1.5
+Gemma4-31B
9.32
23.20
11.79
13.60
Δ Thinking (pp) ↑
+5.1
+5.4
+2.0
+0.8
Table 4: Terminology error rates (%, ↓ ) with thinking. Δ Thinking reports the change in relative error reduction compared with Table 1 (percentage points, ↑ ). Positive values indicate improvement; negative values (bold) indicate degradation.
Automatic speech recognition (ASR) has achieved substantial gains in transcription accuracy, yet verbatim transcription does not necessarily produce readily usable text. It retains fillers, repetitions, false starts, and self-corrections that increase reading effort, obscure the speaker's final intent, and propagate unresolved or abandoned content to downstream tasks. Existing spoken-to-written methods process completed audio or transcripts but cannot revise emitted text when later speech changes how preceding content should be interpreted. We therefore formulate Agentic Speech Recognition (AgenticSR), an audio-to-clean-text task that removes disfluencies, resolves self-corrections, and normalizes written form while preserving the speaker's final intent. AgenticASR implements this task through an ASR--Refiner architecture that repeatedly transforms a bounded active context and replaces its corresponding output span as audio arrives. This enables continual emission and revision over streams of arbitrary duration. We also introduce AASR-Bench, a bilingual benchmark with fine-grained atomic rubrics. Across multiple ASR front ends, AgenticASR attains the highest AASR-Bench scores among evaluated systems. A human--AI agreement study shows that rubric-based judgments align with independent expert assessments. Ablations characterize Refiner capacity, context length, and the quality--latency trade-off between online and offline inference. Together, these results establish AgenticASR as a practical framework for intent-preserving clean transcription during ongoing speech. Code, AASR-Bench, and a demo will be released at https://github.com/AnXMuy/AgenticASR.
Zixuan Jiang, Binghao Qiang, Jiaying Chi +3
X-LANCE lab, Shanghai Jiao Tong University · Shanghai Innovation Institute · College of Artificial Intelligence, Xi’an Jiaotong University
We present Voice Memory, a inference-only scheme for agentic speech recognition: at stream time, a frozen corrector reads a single per-domain memory.md and decides per utterance whether to act on the hypothesis or abstain and keep the 1-best. Asynchronously, a score-gated optimizer revises that file through bounded edits, accepting an edit only when it strictly improves a held-out score. Extended from classical ASR-LM framework, we refer this split the listener-thinker architecture; the two roles are coupled only through the memory, so no weights change and the learned skill stays auditable and portable. Restraint turns out to be the operative skill this loop discovers: unconstrained generative error correction (GER) over-corrects, breaking correct tokens on up to 64% of its edits on financial news, and Voice Memory, reduces this rate to 35%. Across ten HyPoradise domains with an open corrector, Voice Memory, lowers weighted word error rate from 8.36% to 7.52% (7.47% with three added in-context examples) without regressing any dataset below its 1-best baseline; gains concentrate where recoverable headroom is largest, including air-travel commands (8.40% to 3.40%) and noisy far-field speech (CHiME-4, 12.69% to 10.46%). The memory transfers across corrector families and adds zero parameters to the inference path. A demo and example code are provided for future studies.
Chao-Han Huck Yang, Zih-Ching Chen, Piotr Zelasko +3
Recognizing entity phrases remains a critical challenge for speech large language models. Existing prompting methods lack an explicit decoding-time biasing weight, limiting their controllability. Generative error correction methods can introduce hallucinated over-corrections. To address these limitations, we propose LOGIC (logit-space integration for contextual biasing), a robust framework operating directly in the logit space. By decoupling context injection from input processing, LOGIC enables explicit control over the biasing strength. Extensive experiments with an open-source speech large language model across 11 locales demonstrate that LOGIC achieves an average 9% relative reduction in entity word error rate, with an average false alarm rate increase of 0.3% and a 2.8% relative runtime overhead. When combined with prompting, LOGIC can reduce entity word error rate by 5% relative to the prompt-only method.