Natural conversations make both speech recognition and speaker attribution challenging for ASR, as speakers take turns, overlap, and reappear over time. We propose STAM-ASR, Speaker-Temporal Anchoring with Memory, a lightweight framework that extends an already pretrained AudioLLM for multi-speaker ASR. Without relying on an external diarization system, STAM-ASR learns speaker activity and speaker-aware representations directly from intermediate AudioLLM features. Hence providing explicit who and when cues to modulate the AudioLLM's semantic representation without explicit speech separation. STAM-ASR further maintains fixed-size speaker and conversational memories to carry complementary context across turns. We evaluate STAM-ASR on AMI, ICSI, LibriCSS, and NOTSOFAR-1 across close-talk, far-field, overlapping, and cross-domain conditions. Our reported results shows that speaker-temporal conditioning and memory provide complementary benefits, while the gap between reference and predicted speaker activity identifies robust speaker tracking as a key remaining challenge.
Figures & tables
Figure 1: Overview of STAM-ASR with internal diarization, speaker-temporal modulation, memory construction.
Corpus
Sessions
Hours
Spk.
Train Sup.
AMI-IHM
136/18/16
79.1/9.4/8.9
3–5/4/3–4
A,D
ICSI
49/6/5
45.6/5.0/5.4
3–8/5–8/5–10
A,D
Mixer6
189/–/–
51.2/–/–
1/–/–
A
LibriSpeech-Sim.
5.2k/–/–
130/–/–
3–5/–/–
D
AMI-SDM
–/–/16
–/–/8.9
–/–/3–4
A,D
LibriCSS
–/–/60
–/–/10.1
–/–/8
A,D
Table 1: Corpora used for training and evaluation. Train/Dev/Test and A and D represent ASR and diarization supervision (Sup.).
Corpus
Chunks
Sess.
Turn/Ch.
Spk/Ch.
Hours
AMI
14,532
136
4.42
4.00
58.22
ICSI
8,847
49
8.82
6.21
42.80
Mixer6
9,533
189
3.14
1.00
41.84
Total
32,912
374
5.23
3.73
142.86
Table 2: Detailed statistics of training chunk.
Model
AMI-IHM
AMI-SDM
ICSI
LibriCSS
NOTSOFAR-1
Reference-defined segments ( ≈ 20 s)
Whisper-LV3 + pyannote
57.2/58.6 (55.1)
61.6/63.1 (58.7)
58.2/60.7 (56.0)
25.5/26.0 (21.2)
62.9/64.0 (52.1)
Qwen2.5-Omni-7B
58.5/60.0 (57.7)
66.5/68.1 (65.6)
46.9/47.5 (45.6)
21.6/21.8 (21.4)
67.0/67.9 (65.2)
TagSpeech-AMI ⋆
54.3/57.2 (38.2)
55.3/57.8 (40.6)
53.4/55.5 (31.2)
63.2/64.8(32.7)
81.0/83.3(62.3)
STAM-ASR (ref.)
25.3/28.3 (25.0)
38.8/41.5 (38.6)
20.4/21.2 (20.4)
18.5/18.8 (18.5)
52.6/53.7 (52.2)
STAM-ASR (pred.)
51.5/53.0 (44.1)
64.6/66.4 (57.0)
47.7/49.3 (42.5)
53.3/55.9 (35.6)
76.6/79.8/(67.6)
Table 3: Reported cpWER / tcpWER@5 (WER) ( ↓ ) under reference-defined and VAD-based segmentation. Bold marks the best comparable cpWER/tcpWER@5; for STAM-ASR , ref. and pred. means reference and predicted speaker activity, respectively, and italicized results use reference speaker activity.
Inference configuration
AMI-IHM
AMI-SDM
ICSI
LibriCSS
NOTSOFAR-1
Reference speaker activity
All (ST+SM+CM)
25.3 / 28.3 (25.0)
38.8 / 41.5 (38.6)
20.4 / 21.2 (20.4)
18.5 / 18.8 (18.5)
52.6 / 53.7 (52.2)
ST
25.8 / 29.1 (25.5)
33.7 / 37.2 (33.5)
23.7 / 24.5 (23.7)
13.9 / 14.1 (13.9)
47.0 / 48.0 (46.7)
ST+SM
25.2 / 28.3 (24.9)
37.0 / 39.9 (36.8)
20.9 / 21.6 (21.0)
17.1 / 17.4 (17.1)
50.7 / 51.8 (50.2)
ST+CM
25.1 / 28.0 (24.8)
38.4 / 41.0 (38.2)
20.4 / 21.2 (20.3)
19.7 / 20.0 (19.6)
52.1 / 53.2 (51.6)
Predicted speaker activity
Table 4: Inference-time STAM-ASR component analysis using the same jointly trained checkpoint. ST = speaker-temporal conditioning, SM = speaker memory, and CM = conversation memory; All denotes ST+SM+CM. Results are cpWER / tcpWER@5 (WER) on reference-defined segments. Best results are bold; dev sets show consistent trends.
Model
AMI-IHM
AMI-SDM
ICSI
LibriCSS
NSF-1
TagSpeech-AMI
52.1
54.8
44.8
48.9
65.1
STAM-ASR
37.6
49.7
45.0
41.5
65.5
Table 5: Full-session gDI-cpWER (%; ↓ ) using VAD segmentation and scored over complete recordings.
Dϕ
Sortformer
pyannote
TagSpeech ∗
Test
DER
Conf.
DER
Conf.
DER
Conf.
DER
Conf.
AMI-IHM
36.6±3.8 (32.3)
17.4
29.0‡
7.7
12.3
3.7
39.9
7.8
AMI-SDM
51.4±4.8 (49.3)
17.6
35.3‡
13.7
15.4
5.2
37.8
8.3
ICSI
57.0±1.9
28.9
–
–
25.1
3.1
34.7
11.2
LibriCSS
56.7±1.1
45.2
46.7
34.5
11.3
3.7
33.8
15.6
NOTSOFAR-1
40.5±0.8
27.8
22.5
10.2
19.9
11.1
54.2
11.7
Table 6: Diarization error rate (DER) and speaker confusion (Conf.), with a 0.25 s collar. Dϕ reports mean ± s.d. over three seeds. ‡ Sortformer covers 6/16 AMI meetings; parentheses show Dϕ on the same subset. ∗ TagSpeech uses 20–25 s chunk-local scoring, which does not penalize cross-chunk speaker identity.
We present TagSpeech, a unified LLM-based framework that utilizes Temporal Anchor Grounding for joint multi-speaker ASR and diarization. The framework is built on two key designs: (1) decoupled semantic and speaker streams fine-tuned via Serialized Output Training (SOT) to learn turn-taking dynamics; and (2) an interleaved time anchor mechanism that not only supports fine-grained timestamp prediction but also acts as a synchronization signal between semantic understanding and speaker tracking. Compared to previous works that primarily focus on speaker-attributed ASR or implicit diarization, TagSpeech addresses the challenge of fine-grained speaker-content alignment and explicitly models "who spoke what and when" in an end-to-end manner. Experiments on AMI and AliMeeting benchmarks demonstrate that our method achieves consistent improvements in Diarization Error Rate (DER) over strong end-to-end baselines, including Qwen-Omni and Gemini, particularly in handling complex speech overlaps. Moreover, TagSpeech employs a parameter-efficient training paradigm in which the LLM backbone is frozen and only lightweight projectors are trained, resulting in strong performance with low computational cost.
Multi-talker speech recognition is often addressed by combining automatic speech recognition (ASR) and speaker diarization in a pipeline system. Recently, LLM-based approaches have shown promise by jointly modeling semantic and speaker information, but they typically require large-scale multi-talker corpora that are costly to annotate. In this paper, we investigate how to efficiently train an LLM-based system with limited real-recorded data while maintaining high accuracy in speaker attribution. We propose several strategies: (1) a dual-encoder architecture to extract semantic and speaker features, (2) a feature interleaving format to merge these features as the inputs to the LLM, (3) a length-aware speaker ID loss to enhance diarization capability, and (4) an adaptive threshold strategy for ASR loss computation to mitigate hallucinations caused by speech overlaps. These strategies balance training between ASR and diarization tasks. Our system outperforms open-source baseline approaches, achieving relative improvements of 18% on the AliMeeting corpus and 24% on the Aishell4 corpus.
Multi-talker automatic speech recognition (MT-ASR) remains challenging in the presence of overlapping speech. Hard segmentation introduces irreversible errors, whereas serialized output training (SOT) avoids explicit segmentation but does not condition a pretrained encoder on speaker activity. We propose Soft Posterior Speaker Injection (SPSI). A Soft Posterior Head predicts per-frame speaker posteriors P^ and injects them into Whisper through Multi-layer Feature-wise Linear Modulation (MFLM) and Speaker Memory Prompts (SMP). The benefit of SPSI is largest where overlap is heaviest and under domain transfer. On controlled two-speaker LibriSpeech overlap, SPSI reduces concatenated minimum-permutation word error rate (cpWER) from 61.5% to 60.0% in the high-overlap bin, and from 51.9% to 51.0% on the full set, relative to SOT. By contrast, Speaker CE, SD-CTC, SA-DiCoW, and Pipeline (oracle/est.\ VAD) do not outperform SOT. Freeze-posterior overlap-heavy adaptation reduces held-out LibriCSS cpWER from 42.3% to 36.8% on sessions 8--9, a 5.5-point gain over SOT. The source code is available at https://github.com/HackerHyper/SPSI.