eess.ASOct 5, 2026

Logbook: Extremely Long-form Audio Event Understanding

Authors: Kwanghee Choi, Suwon Shon, Dmitriy Serdyuk, Guitang Lan, Chao-Wei Huang, Mohammad Sadegh Rasooli, Sangeeta Srivastava, Zhaojiang Lin, +2 more

Organizations: UT Austin, USA · Meta Reality Labs, USA

Abstract

Audio benchmarks are built around short, pre-segmented clips, limiting model design to brief inputs or fixed vocabularies. To close this gap, we introduce Logbook, a benchmark for hour-scale audio understanding, with recordings ranging from ten minutes to six days. Given a continuous audio recording and an event label vocabulary, a system must predict a gap-free segmentation with an event label and a description per segment. We compare 52 systems, end-to-end and cascaded, and ablate fine-tuning, context length, and reasoning budget. We find the task tractable, though the best systems remain below the human reference. Also, over-segmentation is pervasive, and fine-tuning partially mitigates it. Finally, end-to-end are often better than cascaded systems, but degrades with longer context.

Figures & tables

Appendix figures & tables8 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. LongAudioSpan: Spanning the Duration and Depth of Audio Comprehension

    Aug 26, 2026Wen Huang, Yunfei Chu, Meng Gao +2Audio UnderstandingLarge Audio Language Models

  2. Event-Grounded Question Answering over Long Audio via Structured Retrieval

    Feb 16, 2026Kartik Hegde, Arvind Krishna Sridhar, Naveen Vakada +2Temporal GroundingAudio Understanding

  3. GigaChat Audio: Time-aware Large Audio Language Model

    Jul 11, 2026Aleksandr Kutsakov, Mariia Sadovina, Georgii Gospodinov +4Large Audio Language ModelsText-To-Audio