TEMA: Evidence-Grounded Temporal Question Answering in Multi-Turn Multi-Audio Dialogs
Authors: Kaidi Yang, Hualei Wang, Zhaohui Wang, Chenxuan Wang, Hong Liu, Xiangdong Wang
Organizations: Institute of Computing Technology, Chinese Academy of Sciences, Beijing, China. · University of Chinese Academy of Sciences, Beijing, China.
Multi-turn, multi-audio temporal question answering requires models to track target events across follow-up questions, recording switches, and historical references, recovering complete instances and their boundaries for temporal calculation and comparison. We propose TEMA, which connects event perception with evidence-based answering through Route, specifying the audio scope, and Span, describing all relevant intervals as conditional audio captions. We construct TEMA-Dialog with 40,704 dialogs and per-turn evidence and answer supervision, and TEMA-Bench for joint evaluation of evidence and final answers. Training combines temporal grounding initialization, full-dialog supervised fine-tuning, and completeness-first Span-only GRPO. Experiments on Qwen2.5-Omni and AF-Next show improved temporal question answering, particularly event localization and cross-audio comparison. Reinforcement learning applied solely to evidence further improves interval recovery and answer accuracy.
Figures & tables
Figure 1: Temporal QA and evidence in multi-audio dialog.
Figure 2: Construction pipeline for TEMA-Dialog and TEMA-Bench.
Model
Training
Overall
Evidence
Task-family QA
QA
QA+T
Span F1
H
F1
F2
F3
F4
F5
Qwen2.5-Omni-7B
Original
46.09
38.90
—
—
26.05
76.89
54.84
22.67
33.33
AF-Next-Instruct
Original
51.98
46.57
—
—
39.07
72.79
80.65
29.78
36.11
TEMPO multitask-RL
Original
41.89
19.77
—
—
36.87
62.42
35.48
12.89
33.33
Qwen3.5-Omni-Plus
Original
57.71
46.49
—
—
28.04
89.42
87.10
52.44
5.56
Gemini 3.5 Flash
Original
49.72
39.39
—
—
21.41
80.78
53.23
48.00
11.11
Table 1: Natural multi-turn results and ablations on Test253; gold-evidence diagnostics below. I/S/G: temporal grounding initialization, dialog SFT, and Span-only GRPO. All scores are percentages; family columns report QA. Evidence metrics follow Sec. 3.1 . —: not evaluated. Bold: best available natural score. Percentage-point changes use displayed rounded scores.
Answering natural-language questions over multi-hour audio requires both event recognition and temporal grounding. Current large audio-language models perform well on short clips, but are limited by context length, query-time cost, and weak temporal localization. We present LA-RAG (Long Audio-Retrieval Augmented Generation), a structured framework that converts continuous audio into timestamped event records using an open-vocabulary Audio Grounding Model (AGM), stores them in a SQL event database, and answers queries through intent-aware retrieval followed by LLM-based generation. LA-RAG supports offline grounding mode, where long recordings are pre-indexed for low-latency QA, and inference-time grounding mode, where query-conditioned grounding is performed for shorter open-ended clips. We create 24-hour Home-IoT and Industrial-IoT audio benchmarks and augment CASTELLA, a real-world audio moment retrieval dataset with QA pairs. In offline grounding mode, LA-RAG achieves 76.88% overall accuracy on Home-IoT and 71.10% on Industrial-IoT, with average query latencies below 0.6 seconds. In inference-time grounding mode, state-of-the-art LALMs achieve competitive event-detection accuracy on CASTELLA-QA but low temporal detection F1. We further show that LALMs augmented with our structured retrieval metadata achieve consistent temporal detection improvements, with F1 gains of 11-17% across baseline models with improved latency. These results show that explicit timestamped grounding and structured retrieval provide a practical complement to generative audio-language models for deployment-oriented long-audio QA.
Temporal grounding in long recordings remains challenging for audio-conditioned LLMs. We present a time-aware audio LLM that answers questions with explicit timestamps over up to 120 minutes of input. Our approach interleaves periodic time markers with continuous audio tokens using large-scale synthetic supervision from a cascaded pipeline. Our model achieves strong temporal-grounding accuracy on short and long benchmarks and supports time-anchored fragment descriptions and summaries. Extensive ablations examine how time representation, marker frequency, tokenization, and duration-mixture design affect accuracy and computational cost. We release model weights and datasets to support further research on time-aware audio understanding, available at https://huggingface.co/ai-sage/GigaChat3.1-Audio-10B-A1.8B.
Aleksandr Kutsakov, Mariia Sadovina, Georgii Gospodinov +4
While LALMs show promise on audio question answering, they fail to focus on question-relevant segments of audio and provide a clear, checkable reasoning process when dealing with complex audio reasoning. Reinforcement learning and tool-augmented prompting can help models better relate questions to audio but lack a reliable way to understand, integrate, and self-verify audio segments. To address this gap, we present EChO-Agent, a modular agent framework that reformulates complex audio QA as a planning, tool execution, evidence integration, and answer verification workflow. Experiments on MMAR benchmark show EChO-Agent improves both accuracy and rubric scores over baseline and ablation studies show evidence integration is the key factor.