Revisiting Frame-Wise Saliency for Audio Moment Retrieval
Organizations: LY Corporation, Tokyo, Japan
Abstract
This paper revisits frame-wise saliency for audio moment retrieval (AMR). We show that the frame-wise saliency sequence, conventionally used only as an auxiliary output in DETR-based AMR models, can itself serve as an effective source of moment predictions. We convert the saliency sequence into ranked moments using a simple SED-inspired segmentation rule with no learned parameters, enabling moment retrieval directly from frame-wise temporal information. On the CASTELLA dataset, saliency-based prediction consistently outperforms decoder-based prediction from the same model across all 18 runs of QD-DETR and CG-DETR. For QD-DETR, simply replacing the inference output improves R1@0.7 from 21.0 to 36.1. The advantage remains 7-15 points when the two outputs are evaluated at their independently selected best epochs. The same tendency extends to TaskWeave and UVCOM, whereas TR-DETR shows the opposite behavior, suggesting that how saliency construction may matter. The performance gap is especially pronounced for short moments: for queries whose annotated moments average at most 2 s, R1@0.7 improves from 6.3 to 27.8 with QD-DETR. Decoder supervision nevertheless benefits saliency-based prediction, indicating that its role during training differs from the utility of its inference output.
Figures & tables
| R1@0.5 | R1@0.7 | mAP | |||||
| Model | sal. | dec. | sal. | dec. | sal. | dec. | |
| QD-DETR | 1 | 50.1 1.1 | 40.8 1.6 | 36.1 0.8 | 21.0 1.3 | 30.7 0.5 | 17.5 1.0 |
| 4 | 50.3 0.6 | 39.1 2.5 | 36.1 0.8 | 20.2 2.7 | 30.5 0.2 | 17.4 1.5 | |
| 8 | 51.3 0.5 | 41.9 1.5 | 37.3 1.0 | 21.1 1.3 | 31.4 0.4 | 18.5 0.5 | |
| CG-DETR | 1 | 51.2 1.3 | 46.3 2.9 | 36.7 0.5 | 25.9 3.7 | 29.7 0.5 | 21.3 2.4 |
| 4 | 52.0 1.0 | 48.2 4.3 | 38.3 0.9 | 27.4 5.7 | 30.6 0.4 | 22.4 3.3 | |
| QD-DETR | CG-DETR | |
|---|---|---|
| 1 | 21.4 | 29.8 |
| 4 | 23.2 | 30.2 |
| 8 | 24.3 | 28.2 |
| QD-DETR | CG-DETR | |||
|---|---|---|---|---|
| Peak-relative threshold (Ours) | 36.1 | 37.3 | 36.7 | 38.5 |
| - w/o gap tolerance and overlap removal | 30.2 | 32.1 | 31.3 | 33.4 |
| Recording-wise threshold | 32.5 | 33.0 | 32.7 | 35.4 |
| Global absolute threshold | 21.2 | 24.1 | 25.1 | 24.3 |
| Fixed 5 s duration | 6.1 | 6.2 | 6.0 | 6.5 |
| R1@0.5 | R1@0.7 | mAP | ||||
| Model | sal. | dec. | sal. | dec. | sal. | dec. |
| TaskWeave (HD MR) | 51.9 | 37.6 | 36.5 | 19.2 | 29.9 | 17.0 |
| UVCOM | 50.5 | 47.6 | 35.9 | 28.2 | 30.2 | 24.2 |
| TR-DETR | 41.1 | 52.2 | 22.7 | 32.9 | 18.2 | 26.7 |
| 1 moment | 2–3 | 4 | ||
| Model | Output | ( =573) | ( =404) | ( =370) |
| R1@0.7 | ||||
| QD-DETR ( ) | dec. | 21.8 | 21.3 | 19.3 |
| sal. | 34.3 | 36.1 | 38.9 | |
| CG-DETR ( ) | dec. | 29.7 | 26.7 | 28.5 |
| sal. | 36.9 | 38.4 | 41.0 | |
| 2 s | s | 5 s | ||
|---|---|---|---|---|
| Model | Output | ( =290) | ( =358) | ( =699) |
| QD-DETR ( ) | dec. | 6.3 | 15.7 | 29.7 |
| sal. | 27.8 | 30.7 | 42.3 | |
| CG-DETR ( ) | dec. | 12.3 | 19.5 | 39.7 |
| sal. | 31.6 | 30.7 | 45.4 |
| R1@0.7 | mAP | |||
|---|---|---|---|---|
| Model | joint | no dec. | joint | no dec. |
| QD-DETR | 37.3 1.0 | 34.3 3.2 | 31.4 0.4 | 29.0 1.8 |
| CG-DETR | 38.5 0.6 | 36.4 1.9 | 31.4 0.1 | 29.4 1.4 |