LLM-powered Query Expansion for Enhancing Boundary Prediction in Language-driven Action Localization
Organizations: Beijing Key Laboratory of Intelligent Information Technology, School of Computer Science & Technology, Beijing Institute of Technology, China. · Guangdong Laboratory of Machine Perception and Intelligent Computing, Shenzhen MSU-BIT University, China.
Abstract
Language-driven action localization in videos requires not only semantic alignment between language query and video segment, but also prediction of action boundaries. However, the language query primarily describes the main content of an action and usually lacks specific details of action start and end boundaries, which increases the subjectivity of manual boundary annotation and leads to boundary uncertainty in training data. In this paper, on one hand, we propose to expand the original query by generating textual descriptions of the action start and end boundaries through LLMs, which can provide more detailed boundary cues for localization and thus reduce the impact of boundary uncertainty. On the other hand, to enhance the tolerance to boundary uncertainty during training, we propose to model probability scores of action boundaries by calculating the semantic similarities between frames and the expanded query as well as the temporal distances between frames and the annotated boundary frames. They can provide more consistent boundary supervision, thus improving the stability of training. Our method is model-agnostic and can be seamlessly and easily integrated into any existing models of language-driven action localization in an off-the-shelf manner. Experimental results on several datasets demonstrate the effectiveness of our method.
Figures & tables
| Method | Qvhighlights | Charades-STA | TACoS | ||||||
|---|---|---|---|---|---|---|---|---|---|
| [email protected] | [email protected] | mAP | [email protected] | [email protected] | mAP | [email protected] | [email protected] | mAP | |
| QD-DETR ( Moon et al., 2023b ) | 63.16 | 46.39 | 40.39 | 58.17 | 36.94 | 34.75 | 38.17 | 20.72 | 20.20 |
| QD-DETR+Ours | 64.13 | 47.10 | 41.16 | 59.71 | 38.58 | 35.97 | 40.83 | 24.04 | 22.49 |
| Eatr ( Jang et al., 2023 ) | 58.26 | 41.61 | 36.58 | 55.67 | 33.28 | 33.95 | 31.69 | 16.27 | 16.92 |
| Eatr+Ours | 60.13 | 43.23 | 37.91 | 56.83 | 33.95 | 34.64 | 34.12 | 16.92 | 18.04 |
| TaskWeave ( Yang et al., 2024 ) | 64.26 | 48.17 | 44.20 | 56.85 | 34.44 | 35.37 | 38.52 | 20.82 | 19.99 |
| Module | Charades-STA | QVHighlights | TACoS | |||||||
|---|---|---|---|---|---|---|---|---|---|---|
| QTM | BPM | [email protected] | [email protected] | mAP | [email protected] | [email protected] | mAP | [email protected] | [email protected] | mAP |
| 58.17 | 36.94 | 34.75 | 63.16 | 46.39 | 40.39 | 38.17 | 20.27 | 20.91 | ||
| ✓ | 59.06 | 37.87 | 35.15 | 63.79 | 46.71 | 40.91 | 39.37 | 22.39 | 21.45 | |
| ✓ | 59.33 | 38.25 | 35.41 | 63.60 | 46.82 | 40.57 | 40.04 | 22.14 | 21.75 | |
| ✓ | ✓ | 59.71 | 38.58 | 35.97 | 64.13 | 47.10 | 41.16 | 40.83 | 24.04 | 22.49 |
| LLM | [email protected] | [email protected] | mAP |
|---|---|---|---|
| Noisy query | 59.23 | 38.42 | 35.91 |
| Qwen | 59.51 | 38.10 | 36.13 |
| LLaMa2-7B | 59.46 | 38.28 | 35.81 |
| LLaMa2-13B | 59.30 | 38.31 | 36.22 |
| LLaMa3-8B | 59.71 | 38.58 | 35.97 |
| Prompt Variant | [email protected] | [email protected] | mAP |
|---|---|---|---|
| Prompt-Object | 39.42 | 21.92 | 21.08 |
| Prompt-Concise | 40.61 | 23.99 | 22.76 |
| Ours | 40.83 | 24.04 | 22.49 |
| Charades-STA | TACoS | |||||
|---|---|---|---|---|---|---|
| Probability Modeling | [email protected] | [email protected] | mAP | [email protected] | [email protected] | mAP |
| Gauss distribution | 59.19 | 37.98 | 35.04 | 39.07 | 22.84 | 21.69 |
| Only distance | 59.11 | 37.34 | 35.44 | 40.19 | 23.09 | 21.70 |
| Only similarity | 59.35 | 38.04 | 35.33 | 39.64 | 22.94 | 21.43 |
| Original query | 59.35 | 38.33 | 35.61 | 39.67 | 23.22 | 21.87 |
| Ours | 59.71 | 38.58 | 35.97 | 40.83 | 24.04 | 22.49 |
| Subset | Size | Method | [email protected] | [email protected] | mAP |
|---|---|---|---|---|---|
| Short ( s) | 998 | QD-DETR | 37.68 | 18.64 | 17.37 |
| +Ours | 39.88 | 19.54 | 18.46 | ||
| Average | 2004 | QD-DETR | 44.61 | 25.10 | 23.31 |
| +Ours | 46.86 | 28.99 | 26.37 | ||
| Long ( s) | 999 | QD-DETR | 26.13 | 12.41 | 16.62 |
| +Ours | 30.13 | 14.51 | 18.30 |
| Subset | Size | Method | [email protected] | [email protected] | mAP |
|---|---|---|---|---|---|
| Beginning | 957 | QD-DETR | 70.85 | 48.90 | 44.42 |
| +Ours | 71.16 | 51.93 | 45.25 | ||
| End | 665 | QD-DETR | 62.41 | 46.47 | 42.09 |
| +Ours | 63.46 | 47.37 | 42.49 |
| Process | Time (ms) | FLOPs (G) |
|---|---|---|
| Query Expansion | 1315 | 1442.7 |
| Visual Encoder | 1344 | 54.9 |
| Text Encoder | 2.90 | 1.38 |
| QTM | 19.83 | 3.09 |
| Localization Head | 17.87 | 2.46 |