Audio-visual multi-segment grounding (AV-MSG) in untrimmed videos, reasoning over audio-visual evidence and predicting multiple segments for a query, is a fundamental problem but remains challenging. Visual-only models overlook complementary acoustic cues, while audio-visual models often fail to calibrate the number of events - a phenomenon we refer to as count miscalibration. We present TiTok, an audio-visual large language model (AV-LLM) that localizes an arbitrary number of temporal event segments for each query. For precise boundary prediction, we introduce the Time Token Interleaving (TTI) method, which explicitly injects special time tokens into the audio-visual stream to align input-side temporal perception with output-side temporal prediction. We further propose decoupled, multi-segment-oriented rewards for reinforcement learning, consisting of global, local, count, precision, and format rewards, optimized with Group reward-Decoupled Normalization Policy Optimization (GDPO). To assess the performance on AV-MSG, we establish a new UnAV-100-based evaluation protocol, and propose the CountF1 metric for quantifying count miscalibration that overlap metrics fail to capture. TiTok reaches 65.7 mIoU and 0.58 CountF1, achieving state-of-the-art performance. Our code is available at this link.
Figures & tables
Figure 1 : Dense event captioning vs. audio-visual multi-segment grounding. For the same video, dense event captioning (top) describes every event, whereas audio-visual multi-segment grounding (bottom) localizes only the occurrences of the queried event.
Method
Audio
Video
Multi-Seg
Time Token
ChronusOmni [ 4 ] , ARC-Hunyuan [ 9 ]
✓
✓
✗
✗
TempR1 [ 34 ] , MUSEG [ 22 ]
✗
✓
✓
✗
VTG-LLM [ 12 ]
✗
✓
✗
✓
AVicuna [ 30 ]
✓
✓
✓
✗
TiTok (Ours)
✓
✓
✓
✓
Table 1 : Capability comparison. Multi-Seg marks whether a method natively emits an arbitrary number of segments; Time Token marks whether it represents time with dedicated tokens (special time tokens) rather than plain-text numbers or percentiles.
Figure 2 : The overall pipeline of TiTok. (Left) Cold-start SFT stage. (Right) GDPO-based RL stage with TTI and the five rewards.
Figure 3 : UnAV-100 reconstruction. All intervals of one event category become a single query; a recurring event yields a multi-segment special time token target.
Model
Setting
A
V
mIoU
F1@0.3
F1@0.5
F1@0.7
CountF1
Qwen2.5-Omni †
ZS
✓
✓
12.6
14.2
9.3
6.4
0.07
ARC-Hunyuan †
ZS
✓
✓
45.8
43.6
33.3
26.1
0.00
MUSEG †
ZS
✗
✓
49.1
51.5
36.8
23.0
0.22
ChronusOmni †
ZS
✓
✓
60.9
52.7
34.8
21.8
0.39
AVicuna †
FT
✓
✓
50.4
31.5
23.8
18.6
0.08
Crab + †
FT
✓
✓
50.7
48.4
36.7
27.7
0.24
Table 2 : Results on UnAV-100. Setting column marks each model as ZS (zero-shot) or FT (fine-tuned); A and V mark audio and video support. † marks results we reproduced from released checkpoints.
Figure 4 : Qualitative comparison on UnAV-100. Two queries on the same clip: “playing volleyball” (top) and “people clapping” (bottom).
Setting
mIoU
F1@0.3
F1@0.5
F1@0.7
CountF1
TiTok
65.7
63.8
50.3
38.6
0.58
w/o special tokens
63.9
60.6
47.3
35.5
0.51
w/o interleaving
61.2
59.8
45.6
32.7
0.53
w/o audio
57.8
53.6
40.4
29.4
0.52
w/o RL
56.8
26.9
19.1
14.4
0.16
w/o rprec
65.5
62.9
49.6
36.7
0.59
Table 4 : Ablation on UnAV-100. Each row removes or replaces a component.
Temporal Grounding (TG) aims to localize video segments corresponding to a textual query. Prior research predominantly focuses on single-segment retrieval. Real-world scenarios, however, often require localizing multiple disjoint segments for a single query -- a setting we term One-to-Many Temporal Grounding (OMTG). Previous state-of-the-art MLLMs, optimized for one-to-one settings, struggle in this context, often yielding near-zero scores due to a lack of event cardinality perception. To bridge this gap, we present a systematic solution with three key contributions. First, we establish the first comprehensive OMTG benchmark, introducing Count Accuracy (C-Acc) and Effective Temporal F1 (EtF1) as evaluation metrics. Second, we curate a high-quality OMTG dataset comprising 56k samples through a sophisticated construction pipeline. Third, we develop novel temporal and caption reward functions specifically designed for OMTG. In particular, the caption reward leverages Chain-of-Thought reasoning over dense video captions to explicitly guide policy optimization toward both preciseness and completeness. Extensive experiments show our model achieves a new state-of-the-art EtF1 of 43.65% on OMTG Bench, outperforming Gemini 2.5 Pro and Seed-1.8 by 15.85% and 15.61%, respectively. Project Page: https://insomniaaac.github.io/OMTG/
Qi Xu, Yue Tan, Shihao Chen +5
1Wuhan University · 2Peking University · 3Nanyang Technologi +1
Large Audio Language Models struggle to disentangle overlapping events in complex acoustic scenes, yielding temporally inconsistent captions and frequent hallucinations. We introduce Timestamped Audio Captioner (TAC), a model that produces temporally grounded audio descriptions at varying degrees of detail and resolution. TAC is trained with a synthetic data pipeline that constructs challenging and dynamic mixtures from real-world audio sources, enabling robust learning under realistic polyphonic conditions. Across event detection and dense captioning, TAC outperforms all competing methods, with a low hallucination rate and accurate temporal grounding. We also introduce TAC-V, an audio-visual pipeline to generate semantically rich audio-visual descriptions. We then show that TAC and TAC-V serves as a "semantic bridge" for a text-only reasoner: a simple TAC→LLM and TAC-V→LLM cascade achieves state-of-the-art scores on benchmarks for both audio (MMAU-Pro, MMSU, MMAR) and audio-visual (DailyOmni, VideoHolmes) understanding and reasoning respectively.
Sonal Kumar, Prem Seetharaman, Ke Chen +8
University of Maryland, College Park, USA · Adobe Research, USA · OpenAI, USA (work done while at Adobe).
Video multimodal large language models (MLLMs) can describe what happens in a video, but rarely identify when the supporting evidence occurs. We study generalist video temporal grounding, in which one model predicts a variable-cardinality set of evidence intervals across video lengths, domains, query forms, and viewpoints. Existing training strategies are misaligned with this set-valued task: long-video labels often rely on brittle one-pass annotation, while reinforcement-learning rewards either fail to distinguish non-overlapping predictions or require fragile segment matching. TimeLens2 treats temporal evidence as an interval set throughout supervision and optimization. TimeLens2-93K constructs reliable multi-span supervision through caption-derived proposals, independent localization, cross-agent consensus, semantic verification, and boundary refinement. Our temporal Wasserstein reward computes exact one-dimensional W1 between uniform distributions over merged interval supports, providing dense, matching-free feedback under unequal cardinalities and equivalent fragmentation; temporal IoU complements it with precise-overlap feedback. Across seven benchmarks, TimeLens2-2B outperforms all size-matched baselines on every benchmark, while the 4B and 8B variants achieve state-of-the-art performance, surpassing open-source models with up to 397B parameters. The 2B, 4B, and 8B variants improve over their Qwen3-VL backbones by 14.2, 13.0, and 18.1 mIoU points, respectively.
Yuhan Zhu, Changlian Ma, Xiangyu Zeng +12
Nanjing University · Shanghai AI Laboratory · Shanghai Jiao Tong University +3