Zero-Shot Cue-Grounded Topic Segmentation of Spoken Documents
Organizations: Seoul National University · KAIST
Abstract
Topic segmentation structures spoken documents into coherent sections, facilitating navigation and downstream understanding. The appropriate granularity can vary substantially, ranging from broad thematic shifts to fine-grained subtopics. Existing LLM-based segmenters, however, often struggle to adapt to this variation, causing them to either merge distinct subtopics or over-segment coherent themes. To address this, we introduce Cue-Grounded Segmentation (CGS), a training-free framework that operates without any task-specific supervision. CGS first identifies phrases that explicitly signal the start of a new topic and uses their sentence positions as segment boundaries. When such cues are insufficient, it falls back to semantic segmentation, guided by the document structure inferred during cue extraction. Across six benchmarks and six LLM backbones, CGS consistently outperforms existing baselines, remains robust to noisy ASR transcripts, and achieves these gains with low API cost on proprietary models.
Figures & tables
| Dataset | Length | Description | |
|---|---|---|---|
| YTSeg [ 24 ] | 1,448 | 12.9 | Creator-authored chapters |
| ICSI [ 14 , 6 ] | 75 | 241.0 | Research meetings |
| AMI [ 2 ] | 129 | 66.0 | Multi-party meetings |
| MeetingBank [ 13 , 6 ] | 28 | 72.6 | City-council meetings |
| QMSum [ 28 , 6 ] | 20 | 28.6 | Parliamentary committees |
| SIM [ 6 ] | 100 | 91.6 | Spliced meeting excerpts |
| ( ) | WindowDiff ( ) | Boundary ( ) | ||||||||||||||||||||
| Method | YT | IC | MB | SIM | AMI | QM | Avg. | YT | IC | MB | SIM | AMI | QM | Avg. | YT | IC | MB | SIM | AMI | QM | Avg. | |
| Classical | TextTiling [ 12 ] | .513 | .716 | .648 | .604 | .611 | .630 | .620 | .653 | .996 | .994 | 1.00 | .816 | .999 | .910 | .433 | .036 | .086 | .057 | .121 | .129 | .144 |
| BERT-TT [ 26 ] † | .402 | .473 | .420 | .378 | .420 | .407 | .417 | .409 | .514 | .464 | .434 | .442 | .432 | .449 | .189 | .055 | .133 | .197 | .102 | .188 | .144 | |
| LLM-based | LumberChunker [ 3 ] | .378 | .714 | .641 | .604 | .540 | .623 | .583 | .473 | .983 | .952 | .994 | .732 | .972 | .851 | .539 | .121 | .182 | .119 | .291 | .233 | .247 |
| Mackenzie et al. [ 18 ] | .349 | .533 | .460 | .540 | .459 | .511 | .476 | .421 | .722 | .692 | .844 | .599 | .712 | .665 | .519 | .182 | .293 | .141 | .302 | .226 | .277 | |
| TOC prompt [ 7 ] | .414 | .516 | .472 | .574 | .511 | .495 | .497 | .483 | .632 | .648 | .840 | .606 | .693 | .650 | .453 | .098 | .181 | .090 | .158 | .322 | .217 | |
| Guidance | Saved example |
|---|---|
| Segment count | 3–5 segments |
| Structural phases | Prototype Discussion : Follows the introduction; covers design ergonomics, materials, and cost-saving measures. |
| Boundary signals | Introduction of new documents or evaluation criteria |
| Segment duration | 120–300 seconds |
| SegmentLLM | CGS (ours) | |||||
|---|---|---|---|---|---|---|
| Backbone | ( ) | WD ( ) | ( ) | ( ) | WD ( ) | ( ) |
| Gemini-3.1-FL | .234 | .285 | .539 | .177 | .207 | .557 |
| GPT-5.6-Terra | .328 | .415 | .525 | .181 | .218 | .550 |
| Qwen3.6-27B | .241 | .278 | .495 | .175 | .208 | .543 |
| Gemma-4-26B | .411 | .605 | .363 | .221 | .296 | .486 |
| Qwen3.5-9B | .309 | .334 | .299 | .230 | .279 | .464 |
| Method | Input (k) | Output (k) | Cost | |
|---|---|---|---|---|
| LumberChunker | .583 | 32.36 | 0.825 | 35.59 |
| Mackenzie et al. | .476 | 46.25 | 0.233 | 53.03 |
| TOC prompt | .497 | 14.94 | 0.306 | 17.17 |
| Def-DTS | .340 | 16.37 | 42.221 | 317.48 |
| SegmentLLM | .307 | 14.03 | 2.428 | 26.13 |
| CGS | .208 | 19.63 | 0.603 | 21.87 |
| Configuration | Flash-Lite | Qwen-27B |
|---|---|---|
| CGS | .177 | .175 |
| Stage 1 only ( Forced cues ) | .197 | .187 |
| Stage 2 only | .306 | .285 |
| Cue indices only | .197 | .180 |
| Count-only fallback | .192 | .183 |
| Transcript-only fallback | .249 | .218 |
| Method | Correct | Missed | Incorrect |
|---|---|---|---|
| TextTiling | 61.6 | 38.4 | 1361.8 |
| BERT-TT | 13.3 | 86.7 | 71.6 |
| LumberChunker | 79.3 | 20.7 | 750.6 |
| Mackenzie et al. | 62.4 | 37.6 | 476.2 |
| TOC prompt | 41.4 | 58.6 | 262.4 |
| Def-DTS | 46.4 | 53.6 | 351.0 |
| Corpus | Method | WD | ||
|---|---|---|---|---|
| AMI | SegmentLLM | |||
| CGS | ||||
| ICSI | SegmentLLM | |||
| CGS | ||||
| YTSeg | SegmentLLM | |||
| CGS |