Audio benchmarks are built around short, pre-segmented clips, limiting model design to brief inputs or fixed vocabularies. To close this gap, we introduce Logbook, a benchmark for hour-scale audio understanding, with recordings ranging from ten minutes to six days. Given a continuous audio recording and an event label vocabulary, a system must predict a gap-free segmentation with an event label and a description per segment. We compare 52 systems, end-to-end and cascaded, and ablate fine-tuning, context length, and reasoning budget. We find the task tractable, though the best systems remain below the human reference. Also, over-segmentation is pervasive, and fine-tuning partially mitigates it. Finally, end-to-end are often better than cascaded systems, but degrades with longer context.
Figures & tables
Source
Split
# sess.
Avg. len.
Tot. len.
Label
Domain
# subj.
Ego4D
Train
1350
24.7 min
555.8 h
Inferred
Egocentric
250
Ego4D
Val
247
26.1 min
107.6 h
Inferred
Egocentric
71
Ego4D
Test
163
30.2 min
82.1 h
Inferred
Egocentric
55
EgoLife
Test
170
84.9 min
240.6 h
Inferred
Egocentric
6
SINS
Test
1
148.9 h
148.9 h
Provided
Smart home
1
Table 1: Dataset statistics. # sess. is the number of uninterrupted sessions and # subj. is the number of unique recorded subjects.
Ego4D
SINS
EgoLife
Segmentation
Description
Utility
Segmentation
Segmentation
Description
Model
PFLOPs
bF1 ↑
evER ↓
fAcc ↑
dR ↑
dP ↑
dF1 ↑
Hit ↑
Rddc ↓
bF1 ↑
evER ↓
fAcc ↑
bF1 ↑
evER ↓
fAcc ↑
dR ↑
dP ↑
dF1 ↑
EnCLAP
6.41
0.480
2.528
0.320
0.222
0.164
0.188
0.151
0.349
0.444
2.072
0.287
0.409
2.623
0.308
0.080
0.504
0.138
MSCLAP
0.48
0.487
2.618
0.287
0.219
0.156
0.181
0.142
0.308
0.394
2.661
0.395
0.432
2.606
0.302
0.078
0.468
0.132
Qwen 2.5-o
43.01
0.491
3.519
0.295
0.240
0.146
0.181
0.150
0.282
0.432
2.613
0.395
0.450
4.205
0.288
0.093
0.427
0.150
Qwen 3-o
6.69
0.502
4.444
0.313
0.283
0.140
0.186
0.161
0.282
0.358
3.864
0.354
0.460
4.877
0.295
0.114
0.426
0.178
Table 2: Comparing audio captioning models across three datasets. Metrics averaged across eight text models in § 3.2 . Shaded columns highlight the headline metrics for the segmentation and description tasks. The human row is a proxy rather than a ceiling; see § 2.4 .
Ego4D
SINS
EgoLife
Segmentation
Description
Utility
Segmentation
Segmentation
Description
Model
PFLOPs
bF1 ↑
evER ↓
fAcc ↑
dR ↑
dP ↑
dF1 ↑
Hit ↑
Rddc ↓
bF1 ↑
evER ↓
fAcc ↑
bF1 ↑
evER ↓
fAcc ↑
dR ↑
dP ↑
dF1 ↑
K2-V2
62.11
0.404
1.773
0.117
0.107
0.106
0.101
0.061
0.223
0.244
1.277
0.186
0.322
1.526
0.117
0.035
0.402
0.063
OLMo 3.1
18.77
0.496
3.685
0.179
0.221
0.157
0.184
0.139
0.331
0.332
3.339
0.219
0.498
4.685
0.170
0.100
0.483
0.165
Qwen 2.5
39.60
0.492
3.361
0.367
0.270
0.141
0.183
0.172
0.282
0.423
3.484
0.337
0.436
3.683
0.352
0.099
0.448
0.160
Llama 3.3
41.36
0.500
4.565
0.363
0.272
0.151
0.193
0.162
0.351
0.332
4.688
0.331
0.471
4.724
0.369
0.111
0.457
0.178
Table 3: Comparing text models across three datasets. Metrics averaged across six audio captioning models in § 3.2 .
Ego4D
SINS
EgoLife
Segmentation
Description
Utility
Segmentation
Segmentation
Description
System
PFLOPs
bF1 ↑
evER ↓
fAcc ↑
dR ↑
dP ↑
dF1 ↑
Hit ↑
Rddc ↓
bF1 ↑
evER ↓
fAcc ↑
bF1 ↑
evER ↓
fAcc ↑
dR ↑
dP ↑
dF1 ↑
Casc. Gemini 3.5
—
0.493
2.319
0.397
0.276
0.175
0.215
0.198
0.340
0.612
1.133
0.571
0.429
2.686
0.325
0.087
0.512
0.149
Casc. GPT 5.6
—
0.500
2.863
0.383
0.273
0.182
0.218
0.191
0.340
0.640
1.284
0.527
0.428
3.429
0.327
0.096
0.510
0.162
Casc. Gemma 4
71.41
0.508
2.608
0.398
0.262
0.218
0.238
0.160
0.304
0.547
1.588
0.512
0.430
2.822
0.311
0.085
0.560
0.147
+ SFT
71.12
0.247
0.785
0.520
0.316
0.403
0.355
0.180
0.207
0.573
0.668
0.421
0.254
0.756
0.335
0.044
0.706
0.084
Table 4: Comparing cascaded and E2E systems across three datasets. Cascaded systems use AF 3 as the audio captioning model. Shaded columns highlight the headline metrics for the segmentation and description tasks. The human row is a proxy rather than a ceiling; see § 2.4 .
Figure 1: Visualization of the predicted segmentations for the first 24 hours of SINS dataset using SINS label definition.
Figure 2: Segmentation performance ablation on context window size.
Figure 3: Segmentation performance ablation on reasoning budget.
Appendix figures & tables8 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 4: Description performance ablation on the context window size.
Ego4D
EgoLife
SINS
System
Window
Cnt.
Len.
Cnt.
Len.
Cnt.
Len.
Casc. Gemini 3.5
5min
1.84
162
1.57
191
0.88
341
10min
1.26
236
1.12
267
0.57
525
20min
1.08
275
0.90
330
0.43
695
30min
1.05
282
0.90
332
0.43
697
45min
1.05
283
0.90
328
0.46
659
Appendix
Table 5: Events per 5-minute interval across Ego4D, EgoLife, and SINS datasets. Average counts per five minutes, and average length for each event in seconds.
Figure 5: Description performance ablation on the reasoning budget.
System
Ego4D
EgoLife
Casc. Gemini 3.5
9.17
9.21
Casc. GPT 5.6
18.75
18.94
Casc. Gemma 4
9.89
10.06
+ SFT
4.62
4.91
E2E Gemini 3.5
8.50
10.10
E2E Qwen 3-o
9.72
10.61
Appendix
Table 6: Facts per 5-minute interval for Ego4D and EgoLife datasets. We use first pass annotations for Ego4D to count the facts.
Ego4D
EgoLife
SINS
System
Cnt.
Len.
Cnt.
Len.
Cnt.
Len.
Casc. Gemini 3.5
1.26
236
1.12
267
0.57
525
Casc. GPT 5.6
1.51
195
1.31
223
0.63
474
Casc. Gemma 4
1.40
214
1.17
256
0.73
412
+ SFT
0.30
995
0.23
1296
0.21
1403
E2E Gemini 3.5
0.66
384
0.64
384
0.95
307
Appendix
Table 7: Events per 5-minute interval across Ego4D, EgoLife, and SINS datasets. Average counts per five minutes, and average length for each event in seconds.
System
Budget
Ego4D
EgoLife
SINS
Casc. Gemini 3.5
None
0.6%
0.2%
0.0%
1024
1.7%
3.0%
0.9%
2048
2.4%
2.8%
0.9%
4096
49.2%
48.2%
31.1%
Casc. GPT 5.6
None
0.5%
0.1%
0.0%
Low
1.3%
0.9%
0.4%
Appendix
Table 8: Fraction of predictions not following the required format.
Ego4D
SINS
EgoLife
Segmentation
Description
Utility
Segmentation
Segmentation
Description
Audio Cap.
Text model
bF1 ↑
evER ↓
fAcc ↑
R ↑
P ↑
F1 ↑
Hit ↑
rddc ↓
bF1 ↑
evER ↓
fAcc ↑
bF1 ↑
evER ↓
fAcc ↑
R ↑
P ↑
F1 ↑
EnCLAP
K2-V2
0.445
1.796
0.189
0.156
0.120
0.136
0.087
0.277
0.301
1.364
0.243
0.372
1.829
0.219
0.059
0.482
0.105
OLMo 3.1
0.491
2.919
0.177
0.189
0.150
0.168
0.152
0.344
0.361
2.863
0.172
0.496
3.902
0.212
0.089
0.510
0.151
Qwen 2.5
0.481
2.375
0.374
0.238
0.163
0.194
0.176
0.337
0.497
2.229
0.185
0.381
2.354
0.330
0.078
0.518
0.135
Llama 3.3
0.501
3.830
0.366
0.247
0.149
0.186
0.163
0.383
0.361
3.613
0.247
0.438
3.040
0.376
0.096
0.465
0.160
Appendix
Table 9: Comparing all the combinations of cascaded systems.
Ego4D
Cascaded
End-to-end
Class
Ratio
Gemini 3.5
GPT 5.6
Gemma 4
Gemma 4 + SFT
Gemini 3.5
Qwen 3-o
Qwen 3-o + SFT
Qwen 2.5-o
productive
0.401
0.369
0.336
0.324
0.603
0.503
0.269
0.763
0.945
leisure
0.285
0.601
0.616
0.688
0.695
0.513
0.841
0.680
0.017
food
0.117
0.460
0.408
0.451
0.532
0.447
0.229
0.646
0.000
traveling
0.082
0.157
0.164
0.138
0.054
0.114
0.042
0.211
0.132
other
0.059
0.086
0.133
0.055
0.000
0.191
0.264
0.000
0.000
Appendix
Table 10: Comparison of frame-wise accuracy (fAcc) for each event class across top-line systems. Cascaded systems use Audio Flamingo 3 as the audio captioning model. The Ratio column indicates the proportion of frames belonging to each class within the dataset. Classes are sorted by ratio.
General audio comprehension now covers speech, sound, and music over durations from seconds to hours, driven by large audio-language models (LALMs) that are increasingly omni-modal. Yet the benchmarks that test them still rely on clips of seconds, where scores saturate and models converge; recent long-form efforts extend duration but evaluate long audio much as short clips are. We introduce LongAudioSpan, a benchmark that spans both duration and depth: it pairs audio from 10 minutes to over 2 hours with 3,240 questions across three cognitive levels, namely perception, understanding, and reasoning. Two paths supply the questions, differing in how question content is sourced and how ground truth is obtained. Native QA extracts questions from the audio's content, posing each as a multiple-choice item and an open-ended one graded by detailed rubrics. Anchor QA instead injects ground truth, planting acoustic anchors into the audio and building a perception-to-reasoning chain scored only to the first error. A fully automated pipeline constructs every item through structured captioning, QA generation, and adversarial critic feedback. Evaluating 12 LALMs on LongAudioSpan, we find the hard part comes before reasoning: distilling a few relevant facts from a long, redundant signal. This difficulty grows with audio length and falls hardest on perception, especially temporal grounding. LongAudioSpan is available at https://huggingface.co/datasets/holvan/LongAudioSpan.
Wen Huang, Yunfei Chu, Meng Gao +2
Qwen Team, Alibaba Group · Tsinghua University · The Chinese University of Hong Kong
Answering natural-language questions over multi-hour audio requires both event recognition and temporal grounding. Current large audio-language models perform well on short clips, but are limited by context length, query-time cost, and weak temporal localization. We present LA-RAG (Long Audio-Retrieval Augmented Generation), a structured framework that converts continuous audio into timestamped event records using an open-vocabulary Audio Grounding Model (AGM), stores them in a SQL event database, and answers queries through intent-aware retrieval followed by LLM-based generation. LA-RAG supports offline grounding mode, where long recordings are pre-indexed for low-latency QA, and inference-time grounding mode, where query-conditioned grounding is performed for shorter open-ended clips. We create 24-hour Home-IoT and Industrial-IoT audio benchmarks and augment CASTELLA, a real-world audio moment retrieval dataset with QA pairs. In offline grounding mode, LA-RAG achieves 76.88% overall accuracy on Home-IoT and 71.10% on Industrial-IoT, with average query latencies below 0.6 seconds. In inference-time grounding mode, state-of-the-art LALMs achieve competitive event-detection accuracy on CASTELLA-QA but low temporal detection F1. We further show that LALMs augmented with our structured retrieval metadata achieve consistent temporal detection improvements, with F1 gains of 11-17% across baseline models with improved latency. These results show that explicit timestamped grounding and structured retrieval provide a practical complement to generative audio-language models for deployment-oriented long-audio QA.
Temporal grounding in long recordings remains challenging for audio-conditioned LLMs. We present a time-aware audio LLM that answers questions with explicit timestamps over up to 120 minutes of input. Our approach interleaves periodic time markers with continuous audio tokens using large-scale synthetic supervision from a cascaded pipeline. Our model achieves strong temporal-grounding accuracy on short and long benchmarks and supports time-anchored fragment descriptions and summaries. Extensive ablations examine how time representation, marker frequency, tokenization, and duration-mixture design affect accuracy and computational cost. We release model weights and datasets to support further research on time-aware audio understanding, available at https://huggingface.co/ai-sage/GigaChat3.1-Audio-10B-A1.8B.
Aleksandr Kutsakov, Mariia Sadovina, Georgii Gospodinov +4