Natural conversations make both speech recognition and speaker attribution challenging for ASR, as speakers take turns, overlap, and reappear over time. We propose STAM-ASR, Speaker-Temporal Anchoring with Memory, a lightweight framework that extends an already pretrained AudioLLM for multi-speaker ASR. Without relying on an external diarization system, STAM-ASR learns speaker activity and speaker-aware representations directly from intermediate AudioLLM features. Hence providing explicit who and when cues to modulate the AudioLLM's semantic representation without explicit speech separation. STAM-ASR further maintains fixed-size speaker and conversational memories to carry complementary context across turns. We evaluate STAM-ASR on AMI, ICSI, LibriCSS, and NOTSOFAR-1 across close-talk, far-field, overlapping, and cross-domain conditions. Our reported results shows that speaker-temporal conditioning and memory provide complementary benefits, while the gap between reference and predicted speaker activity identifies robust speaker tracking as a key remaining challenge.
Figures & tables
Figure 1: Overview of STAM-ASR with internal diarization, speaker-temporal modulation, memory construction.
Corpus
Sessions
Hours
Spk.
Train Sup.
AMI-IHM
136/18/16
79.1/9.4/8.9
3–5/4/3–4
A,D
ICSI
49/6/5
45.6/5.0/5.4
3–8/5–8/5–10
A,D
Mixer6
189/–/–
51.2/–/–
1/–/–
A
LibriSpeech-Sim.
5.2k/–/–
130/–/–
3–5/–/–
D
AMI-SDM
–/–/16
–/–/8.9
–/–/3–4
A,D
LibriCSS
–/–/60
–/–/10.1
–/–/8
A,D
Table 1: Corpora used for training and evaluation. Train/Dev/Test and A and D represent ASR and diarization supervision (Sup.).
Corpus
Chunks
Sess.
Turn/Ch.
Spk/Ch.
Hours
AMI
14,532
136
4.42
4.00
58.22
ICSI
8,847
49
8.82
6.21
42.80
Mixer6
9,533
189
3.14
1.00
41.84
Total
32,912
374
5.23
3.73
142.86
Table 2: Detailed statistics of training chunk.
Model
AMI-IHM
AMI-SDM
ICSI
LibriCSS
NOTSOFAR-1
Reference-defined segments ( ≈ 20 s)
Whisper-LV3 + pyannote
57.2/58.6 (55.1)
61.6/63.1 (58.7)
58.2/60.7 (56.0)
25.5/26.0 (21.2)
62.9/64.0 (52.1)
Qwen2.5-Omni-7B
58.5/60.0 (57.7)
66.5/68.1 (65.6)
46.9/47.5 (45.6)
21.6/21.8 (21.4)
67.0/67.9 (65.2)
TagSpeech-AMI ⋆
54.3/57.2 (38.2)
55.3/57.8 (40.6)
53.4/55.5 (31.2)
63.2/64.8(32.7)
81.0/83.3(62.3)
STAM-ASR (ref.)
25.3/28.3 (25.0)
38.8/41.5 (38.6)
20.4/21.2 (20.4)
18.5/18.8 (18.5)
52.6/53.7 (52.2)
STAM-ASR (pred.)
51.5/53.0 (44.1)
64.6/66.4 (57.0)
47.7/49.3 (42.5)
53.3/55.9 (35.6)
76.6/79.8/(67.6)
Table 3: Reported cpWER / tcpWER@5 (WER) ( ↓ ) under reference-defined and VAD-based segmentation. Bold marks the best comparable cpWER/tcpWER@5; for STAM-ASR , ref. and pred. means reference and predicted speaker activity, respectively, and italicized results use reference speaker activity.
Inference configuration
AMI-IHM
AMI-SDM
ICSI
LibriCSS
NOTSOFAR-1
Reference speaker activity
All (ST+SM+CM)
25.3 / 28.3 (25.0)
38.8 / 41.5 (38.6)
20.4 / 21.2 (20.4)
18.5 / 18.8 (18.5)
52.6 / 53.7 (52.2)
ST
25.8 / 29.1 (25.5)
33.7 / 37.2 (33.5)
23.7 / 24.5 (23.7)
13.9 / 14.1 (13.9)
47.0 / 48.0 (46.7)
ST+SM
25.2 / 28.3 (24.9)
37.0 / 39.9 (36.8)
20.9 / 21.6 (21.0)
17.1 / 17.4 (17.1)
50.7 / 51.8 (50.2)
ST+CM
25.1 / 28.0 (24.8)
38.4 / 41.0 (38.2)
20.4 / 21.2 (20.3)
19.7 / 20.0 (19.6)
52.1 / 53.2 (51.6)
Predicted speaker activity
Table 4: Inference-time STAM-ASR component analysis using the same jointly trained checkpoint. ST = speaker-temporal conditioning, SM = speaker memory, and CM = conversation memory; All denotes ST+SM+CM. Results are cpWER / tcpWER@5 (WER) on reference-defined segments. Best results are bold; dev sets show consistent trends.
Model
AMI-IHM
AMI-SDM
ICSI
LibriCSS
NSF-1
TagSpeech-AMI
52.1
54.8
44.8
48.9
65.1
STAM-ASR
37.6
49.7
45.0
41.5
65.5
Table 5: Full-session gDI-cpWER (%; ↓ ) using VAD segmentation and scored over complete recordings.
Dϕ
Sortformer
pyannote
TagSpeech ∗
Test
DER
Conf.
DER
Conf.
DER
Conf.
DER
Conf.
AMI-IHM
36.6±3.8 (32.3)
17.4
29.0‡
7.7
12.3
3.7
39.9
7.8
AMI-SDM
51.4±4.8 (49.3)
17.6
35.3‡
13.7
15.4
5.2
37.8
8.3
ICSI
57.0±1.9
28.9
–
–
25.1
3.1
34.7
11.2
LibriCSS
56.7±1.1
45.2
46.7
34.5
11.3
3.7
33.8
15.6
NOTSOFAR-1
40.5±0.8
27.8
22.5
10.2
19.9
11.1
54.2
11.7
Table 6: Diarization error rate (DER) and speaker confusion (Conf.), with a 0.25 s collar. Dϕ reports mean ± s.d. over three seeds. ‡ Sortformer covers 6/16 AMI meetings; parentheses show Dϕ on the same subset. ∗ TagSpeech uses 20–25 s chunk-local scoring, which does not penalize cross-chunk speaker identity.