CoEvoWhen: Policy-Tool Coevolution for Ultra-Long Video Temporal Grounding
Organizations: Zhejiang University, State Key Lab of CAD & CG
Abstract
Ultra-long video temporal grounding requires balancing long-range evidence search with fine-grained event understanding under a limited visual budget, yet existing agentic methods still rely largely on predefined policies and tool capabilities. Motivated by this, we propose a novel policy-tool coevolution framework that jointly evolves high-level policies and executable media tools from the agentic reasoning trajectories of a VLM, forming a reusable skill without updating model parameters. During evolution, an external skill updater distills transferable task experience in long-video temporal grounding, accordingly refining the orchestration of long-range image-based and fine-grained video-based observations. Alongside these policy updates, the updater employs its coding capabilities to upgrade existing tools or create new ones, adapting the tools to long-video evidence acquisition. Equipped with the evolved skill, the VLM autonomously orchestrates tools under the guidance of the evolved policy, coordinating image and video observations for agentic inference without relying on a separate, stronger planning model. Extensive experiments spanning five benchmarks and three VLMs show that policy-tool coevolution consistently improves temporal grounding accuracy in ultra-long videos while reducing visual token cost at inference, and that the evolved skill yields substantial performance gains on general long-video QA without additional task-specific evolution, demonstrating the effectiveness and generalizability of our framework for long-video understanding.
Figures & tables
| VUE-LVTR | ExtremeWhenBench | CoMET-Bench | |||||||||
| 30.1–105.5 min [-1pt](mean 50.9 min) | 45.0–542.5 min [-1pt](mean 75.8 min) | 30.0–123.7 min [-1pt](mean 50.5 min) | |||||||||
| Method | Setting | Precision [-1pt] AUC | Recall [-1pt] AUC | IoU [-1pt] AUC | IoU [-1pt] @0.5 | Mean [-1pt] IoU | Recall [-1pt] @0.5 | Mean [-1pt] IoU | Recall [-1pt] @0.5 | F1 [-1pt] @0.5 | Rejection [-1pt] F1 |
| Qwen3.5-27B | 768 frames | 0.2645 | 0.2657 | 0.1904 | 0.1889 | 0.0392 | 0.0233 | 0.0683 | 0.0708 | 0.0455 | 55.05 |
| InternVL3.5-8B | 128 frames | 0.1452 | 0.1649 | 0.0804 | 0.0619 | 0.0048 | 0.0018 | 0.0098 | 0.0037 | 0.0018 | 57.18 |
| TimeLens-7B | 384 frames | 0.3665 | 0.3873 | 0.2758 | 0.2866 | 0.1024 | 0.0612 | 0.0277 | 0.0129 | 0.0126 | 72.17 |
| Gemini 2.5 Flash | 128 frames | 0.1670 | 0.1815 | 0.0907 | 0.0684 | 0.0089 | 0.0057 | 0.0498 | 0.0369 | 0.0088 | 57.41 |
| Token Efficiency -1pt | Interaction Efficiency -1pt | ||||||||
|---|---|---|---|---|---|---|---|---|---|
| Benchmark | Method | Visual [-1pt] Tokens | Image [-1pt] Tokens | Video [-1pt] Tokens | Model [-1pt] Calls | Local Tool [-1pt] Calls | Visual [-1pt] Obs. | Image [-1pt] Obs. | Video [-1pt] Obs. |
| VUE-LVTR [2pt]30.1–105.5 min [-1pt](mean 50.9 min) | Qwen3.5-27B + Base Skill | 202.58 | 162.01 | 40.57 | 17.03 | 8.16 | 8.11 | 7.35 | 0.76 |
| Qwen3.5-27B + Evolved Skill | 141.89 | 141.37 | 0.52 | 9.63 | 4.79 | 3.86 | 3.82 | 0.04 | |
| (Evolved Base) | 60.69 | 20.64 | 40.05 | 7.40 | 3.37 | 4.25 | 3.53 | 0.72 | |
| 29.96% | 12.74% | 98.72% | 43.5% | 41.3% | 52.4% | 48.0% | 94.7% | ||
Appendix figures & tables12 assets
Supplementary material from the paper’s appendix.
Appendix
| Property | Value |
|---|---|
| Video Duration (Range / Mean) | 30.35–118.61 min / 52.90 min |
| Keyword / Phrase / Sentence | 18 / 34 / 48 |
| Single-Interval / Multi-Interval | 56 / 44 |
| Sparse-Target Queries | 80 |
| Boundary-Sensitive Queries | 88 |
| Video-Edge Queries | 30 |
| Tool | Role | Function |
|---|---|---|
| probe_media | Media Preparation | Read metadata such as video duration, resolution, and frame rate |
| extract_frames_at | Media Preparation | Extract frames at specified source-video timestamps |
| extract_video_clip | Media Preparation | Extract a video clip without audio from a specified time window |
| inspect_images | Image Observation | Add images to the VLM’s visual context |
| inspect_video | Video Observation | Add video to the VLM’s visual context |
| final_answer | Prediction | Submit the final answer in the format required by the current task |
| Benchmark | Total queries | Evolution queries | Evaluation queries |
|---|---|---|---|
| VUE-LVTR | 407 | 100 | 307 |
| ExtremeWhenBench | 2,273 | 0 | 2,273 |
| CoMET-Bench | 2,789 | 0 | 1,599 |
| Parameter | Setting |
|---|---|
| Serving Engine | vLLM |
| Numerical Precision | bfloat16 |
| Maximum Context Length | 262,144 tokens |
| Thinking Mode | Enabled |
| Decoding | Greedy |
| Temperature / Top-p / Top-k | 0 / 1.0 / 0 |
| VLM | Method | Precision AUC | Recall AUC | IoU AUC | IoU@0.3 | IoU@0.5 | IoU@0.7 |
|---|---|---|---|---|---|---|---|
| Qwen3.5-9B | + Base Skill | 0.1759 | 0.1961 | 0.1142 | 0.1600 | 0.1300 | 0.0500 |
| + Evolved Skill | 0.3065 | 0.3545 | 0.2024 | 0.3000 | 0.2000 | 0.1100 | |
| Evolution Gain | +0.1306 | +0.1584 | +0.0882 | +0.1400 | +0.0700 | +0.0600 | |
| 74.2% | 80.8% | 77.2% | 87.5% | 53.8% | 120.0% | ||
| Qwen3.5-27B | + Base Skill | 0.3620 | 0.3047 | 0.2556 | 0.3400 | 0.2800 | 0.1700 |
| + Evolved Skill | 0.4285 | 0.3916 | 0.3207 | 0.4300 | 0.3500 | 0.2400 |
| Benchmark | Method | Visual Tokens | Image Tokens | Video Tokens |
|---|---|---|---|---|
| VUE-LVTR held-out | Base Skill | 241.21 | 214.98 | 26.23 |
| Evolved Skill | 139.36 | 134.10 | 5.26 | |
| (Evolved Base) | 101.85 | 80.88 | 20.97 | |
| 42.2% | 37.6% | 79.9% | ||
| ExtremeWhenBench | Base Skill | 220.29 | 195.24 | 25.05 |
| Evolved Skill | 158.33 | 152.64 | 5.69 |
| VLM | Method | Action | Environment | Object | Reaction | Scene |
|---|---|---|---|---|---|---|
| Qwen3.5-9B | + Base Skill | 0.0543 | 0.0986 | 0.0877 | 0.0439 | 0.0570 |
| + Evolved Skill | 0.1092 | 0.1485 | 0.1518 | 0.1218 | 0.1246 | |
| Evolution Gain | +0.0549 | +0.0499 | +0.0641 | +0.0779 | +0.0676 | |
| 101.1% | 50.6% | 73.1% | 177.4% | 118.6% | ||
| Qwen3.5-27B | + Base Skill | 0.1446 | 0.2040 | 0.1862 | 0.1167 | 0.1377 |
| + Evolved Skill | 0.2459 | 0.2884 | 0.3427 | 0.1872 | 0.2664 |
| VLM | Method | MAE | OBO | Pearson |
|---|---|---|---|---|
| Qwen3.5-9B | + Base Skill | 3.7104 | 0.5378 | 0.3305 |
| + Evolved Skill | 3.6929 | 0.5516 | 0.6614 | |
| Evolution Gain | 0.0175 | +0.0138 | +0.3309 | |
| 0.5% | 2.6% | 100.1% | ||
| Qwen3.5-27B | + Base Skill | 3.6717 | 0.5760 | 0.0775 |
| + Evolved Skill | 3.4290 | 0.6116 | 0.1258 |
| VLM | Method | Rejection-F1 | FPR | PosCoverage |
|---|---|---|---|---|
| Qwen3.5-9B | + Base Skill | 67.51 | 0.0989 | 0.5397 |
| + Evolved Skill | 68.80 | 0.0968 | 0.5556 | |
| Evolution Gain | +1.29 | 0.0021 | +0.0159 | |
| 1.9% | 2.1% | 2.9% | ||
| Qwen3.5-27B | + Base Skill | 66.13 | 0.1183 | 0.5291 |
| + Evolved Skill | 77.06 | 0.1032 | 0.6755 |
| VLM | Method | 1 event | 2–3 | 4–7 | 8–15 | 16 |
|---|---|---|---|---|---|---|
| Qwen3.5-9B | + Base Skill | 0.1055 | 0.0673 | 0.0460 | 0.0376 | 0.0124 |
| + Evolved Skill | 0.1314 | 0.1296 | 0.0574 | 0.0622 | 0.0239 | |
| Evolution Gain | +0.0259 | +0.0623 | +0.0114 | +0.0246 | +0.0115 | |
| 24.5% | 92.6% | 24.8% | 65.4% | 92.7% | ||
| Qwen3.5-27B | + Base Skill | 0.2058 | 0.1846 | 0.1257 | 0.0860 | 0.0226 |
| + Evolved Skill | 0.2615 | 0.2101 | 0.1498 | 0.0939 | 0.0529 |