cs.CVSep 30, 2026

CoEvoWhen: Policy-Tool Coevolution for Ultra-Long Video Temporal Grounding

Authors: Yiduo Jia, Muzhi Zhu, Jinchuan Shi, Hao Zhong, Yuling Xi, Ke Liu, Hao Chen

Organizations: Zhejiang University, State Key Lab of CAD & CG

Abstract

Ultra-long video temporal grounding requires balancing long-range evidence search with fine-grained event understanding under a limited visual budget, yet existing agentic methods still rely largely on predefined policies and tool capabilities. Motivated by this, we propose a novel policy-tool coevolution framework that jointly evolves high-level policies and executable media tools from the agentic reasoning trajectories of a VLM, forming a reusable skill without updating model parameters. During evolution, an external skill updater distills transferable task experience in long-video temporal grounding, accordingly refining the orchestration of long-range image-based and fine-grained video-based observations. Alongside these policy updates, the updater employs its coding capabilities to upgrade existing tools or create new ones, adapting the tools to long-video evidence acquisition. Equipped with the evolved skill, the VLM autonomously orchestrates tools under the guidance of the evolved policy, coordinating image and video observations for agentic inference without relying on a separate, stronger planning model. Extensive experiments spanning five benchmarks and three VLMs show that policy-tool coevolution consistently improves temporal grounding accuracy in ultra-long videos while reducing visual token cost at inference, and that the evolved skill yields substantial performance gains on general long-video QA without additional task-specific evolution, demonstrating the effectiveness and generalizability of our framework for long-video understanding.

Figures & tables

Appendix figures & tables12 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. VideoEvolve: Co-Evolving Memory and Retrieval for Long Video Understanding

    Oct 7, 2026Yongchao Xu, Bowen Ye, Jiefeng Gan +5Long-Video UnderstandingMemory-Augmented Video Understanding

  2. MetaVideoAgent: Automated Video-Agent Evolution for Long-Form Video Understanding

    Aug 5, 2026Benlei Cui, Ruize Wang, Junjie Li +9Video QALong-Video Understanding

  3. VideoEvolve: Evolving Agent Harnesses for Video Temporal Grounding

    Oct 1, 2026Bingjun Luo, Yuhuan Fan, Jialin Guo +1Temporal Video GroundingEvolutionary Optimization