Topic segmentation structures spoken documents into coherent sections, facilitating navigation and downstream understanding. The appropriate granularity can vary substantially, ranging from broad thematic shifts to fine-grained subtopics. Existing LLM-based segmenters, however, often struggle to adapt to this variation, causing them to either merge distinct subtopics or over-segment coherent themes. To address this, we introduce Cue-Grounded Segmentation (CGS), a training-free framework that operates without any task-specific supervision. CGS first identifies phrases that explicitly signal the start of a new topic and uses their sentence positions as segment boundaries. When such cues are insufficient, it falls back to semantic segmentation, guided by the document structure inferred during cue extraction. Across six benchmarks and six LLM backbones, CGS consistently outperforms existing baselines, remains robust to noisy ASR transcripts, and achieves these gains with low API cost on proprietary models.
Figures & tables
Dataset
n
Length
Description
YTSeg [ 24 ]
1,448
12.9
Creator-authored chapters
ICSI [ 14 , 6 ]
75
241.0
Research meetings
AMI [ 2 ]
129
66.0
Multi-party meetings
MeetingBank [ 13 , 6 ]
28
72.6
City-council meetings
QMSum [ 28 , 6 ]
20
28.6
Parliamentary committees
SIM [ 6 ]
100
91.6
Spliced meeting excerpts
Table 1: Evaluated corpus subsets. n counts documents. For each document, we calculate the average ground-truth segment length in sentences. Length reports the median across documents.
Pk ( ↓ )
WindowDiff ( ↓ )
Boundary F1 ( ↑ )
Method
YT
IC
MB
SIM
AMI
QM
Avg.
YT
IC
MB
SIM
AMI
QM
Avg.
YT
IC
MB
SIM
AMI
QM
Avg.
Classical
TextTiling [ 12 ]
.513
.716
.648
.604
.611
.630
.620
.653
.996
.994
1.00
.816
.999
.910
.433
.036
.086
.057
.121
.129
.144
BERT-TT [ 26 ] †
.402
.473
.420
.378
.420
.407
.417
.409
.514
.464
.434
.442
.432
.449
.189
.055
.133
.197
.102
.188
.144
LLM-based
LumberChunker [ 3 ]
.378
.714
.641
.604
.540
.623
.583
.473
.983
.952
.994
.732
.972
.851
.539
.121
.182
.119
.291
.233
.247
Mackenzie et al. [ 18 ]
.349
.533
.460
.540
.459
.511
.476
.421
.722
.692
.844
.599
.712
.665
.519
.182
.293
.141
.302
.226
.277
TOC prompt [ 7 ]
.414
.516
.472
.574
.511
.495
.497
.483
.632
.648
.840
.606
.693
.650
.453
.098
.181
.090
.158
.322
.217
Table 2: Performance on the six benchmark corpora. Bold and underline mark the best and second-best values per column. † Training-free, but uses self-supervised BERT pretraining. ‡ Def-DTS uses only valid outputs completed within the output-token limit of backbones.
Guidance
Saved example
Segment count
3–5 segments
Structural phases
Prototype Discussion : Follows the introduction; covers design ergonomics, materials, and cost-saving measures.
Boundary signals
Introduction of new documents or evaluation criteria
Segment duration
120–300 seconds
Table 3: Stage-1 guidance supplied to Stage 2: an example from an AMI meeting. Phase and signal entries are excerpts; both ranges are shown in full.
SegmentLLM
CGS (ours)
Backbone
Pk ( ↓ )
WD ( ↓ )
F1 ( ↑ )
Pk ( ↓ )
WD ( ↓ )
F1 ( ↑ )
Gemini-3.1-FL
.234
.285
.539
.177
.207
.557
GPT-5.6-Terra
.328
.415
.525
.181
.218
.550
Qwen3.6-27B
.241
.278
.495
.175
.208
.543
Gemma-4-26B
.411
.605
.363
.221
.296
.486
Qwen3.5-9B
.309
.334
.299
.230
.279
.464
Table 4: Performance by backbone, averaged equally over the six corpora. Table 2 averages LLM results over these backbones. Bold marks the better value.
Method
Pk↓
Input (k) ↓
Output (k) ↓
Cost ↓
LumberChunker
.583
32.36
0.825
35.59
Mackenzie et al.
.476
46.25
0.233
53.03
TOC prompt
.497
14.94
0.306
17.17
Def-DTS
.340
16.37
42.221
317.48
SegmentLLM
.307
14.03
2.428
26.13
CGS
.208
19.63
0.603
21.87
Table 5: Pk from Table 2 and recorded token use (thousands per document). Per-document means are averaged across the six corpora and six backbones. Costs are calculated from recorded token usage at standard API rates and averaged over the two proprietary backbones (USD/1,000 documents). Bold and underline mark the best and second-best values.
Configuration
Flash-Lite
Qwen-27B
CGS
.177
.175
Stage 1 only ( Forced cues )
.197
.187
Stage 2 only
.306
.285
Cue indices only
.197
.180
Count-only fallback
.192
.183
Transcript-only fallback
.249
.218
Table 6: CGS ablations ( Pk↓ , averaged equally over six corpora; T=.3 ). Bold marks the lowest displayed value.
Method
Correct ↑
Missed ↓
Incorrect ↓
TextTiling
61.6
38.4
1361.8
BERT-TT
13.3
86.7
71.6
LumberChunker
79.3
20.7
750.6
Mackenzie et al.
62.4
37.6
476.2
TOC prompt
41.4
58.6
262.4
Def-DTS
46.4
53.6
351.0
Table 7: Counts are normalized to 100 ground-truth boundaries. Incorrect predictions can exceed 100 when a method predicts too many boundaries. Correct counts predictions that match a ground-truth boundary within two sentences. Each boundary is matched at most once. Missed reports how many ground-truth boundaries are not detected. Incorrect counts wrong boundary predictions. Values average six corpora and, for LLMs, six backbones.
Corpus
Method
Pk↓
WD ↓
F1↑
AMI
SegmentLLM
.278
.359
.420
CGS
.202
.263
.512
ICSI
SegmentLLM
.310
.391
.355
CGS
.244
.307
.456
YTSeg
SegmentLLM
.277
.335
.577
CGS
.253
.286
.592
Table 8: ASR segmentation performance, averaged over six backbones. Sentence divisions differ between the original and ASR transcripts.
Dialogue topic segmentation is critical in many human-AI collaborative applications which requires identifying heterogeneous boundary cues, including lexical transitions near utterance edges and semantic discontinuities across utterances. Existing utterance models often dilute these local lexical signals. We propose CobSeg, a novel multi-branch architecture that separates coherence-level semantic continuity from lexical boundary transitions and recovers both through directional boundary prediction. CobSeg further uses boundary informativeness weighting to emphasize high-utility utterance positions, and incorporates a corpus-derived topic coherence cue with learned combination weights. While CobSeg is evaluated as a compact trainable segmenter under supervised gold-boundary training and a pseudo-label setting with automatically induced boundaries, it performs enhanced boundary prediction without LLM calls during inference. Across five benchmarks, it improves Pk and Wd particularly when local lexical cues are prominent: under gold supervision, it reduces Pk by 0.7 points and Wd by 0.6 points on VHF, and reaches Pk of 1.0 on DialSeg711; with induced boundaries, it reduces Pk by 14.8 points on VHF, by 1.5 points on DialSeg711, and by 1.1 points on TIAGE, outperforming prior non-LLM approaches.
Sijin Sun, Liangbin Zhao, Jiaxiang Cai +3
Institute of High Performance Computing, Agency for Science, Technology and Technology · Shanghai Univeristy · Fudan University
While Large Language Models excel in natural language processing, efficiently extending their capabilities to spoken input remains a significant challenge. Existing methods for building SpeechLLMs often rely on computationally expensive full-model fine-tuning, or employ parameter-efficient projectors that suffer from inefficient token sequence lengths and costly full-model supervision. In this paper, we introduce Aligned Continuous Integrate-and-Fire, a highly efficient framework for zero-shot speech processing. Our method dynamically compresses continuous acoustic frames into the exact discrete token length of the target text utilizing explicit Dynamic Time Warping alignments. This allows our initial training stage to establish a robust acoustic-to-semantic bridge using lightweight distance metrics, entirely bypassing the computationally expensive LLM forward pass. For subsequent fine-tuning, we propose a memory-efficient knowledge distillation objective that targets a single LLM layer, performing competitively with full-model cross-entropy training at a fraction of the computational cost. Through extensive evaluations on Automatic Speech Recognition and Speech Translation, we demonstrate that our method achieves superior performance compared to prior parameter-efficient baselines.
Abderrahmane Issam, Yusuf Can Semerci, Jan Scholtes +1
Department of Advanced Computing Sciences Maastricht University
Large-scale pre-trained audio and image models demonstrate an unprecedented degree of generalization, making them suitable for a wide range of applications. Here, we tackle the specific task of sound-prompted segmentation, aiming to segment image regions corresponding to objects heard in an audio signal. Most existing approaches tackle this problem by fine-tuning pre-trained models or by training additional modules specifically for the task. We adopt a different strategy: we introduce a training-free approach that leverages Non-negative Matrix Factorization (NMF) to co-factorize audio and visual features from pre-trained models so as to reveal shared interpretable concepts. These concepts are passed on to an open-vocabulary segmentation model for precise segmentation maps. By using frozen pre-trained models, our method achieves high generalization and establishes state-of-the-art performance in unsupervised sound-prompted segmentation, significantly surpassing previous unsupervised methods.
Hugo Malard, Michel Olvera, Stephane Lathuiliere +1
LTCI, Télécom Paris, Institut Polytechnique de Paris