cs.CVSep 29, 2026

VidHarness: Evolving Agent Harnesses for Cost-Efficient Long Video Understanding

Authors: Susan Liang, Jianmin Wu, Daxiang Dong

Organizations: Baidu, Inc.

Abstract

Vision-language models (VLMs) can answer questions about hour-long videos, but processing every frame is prohibitively expensive, even though the evidence for a question usually spans only a few seconds. Video agents, i.e., harness programs wrapped around a frozen VLM, address this by observing the video selectively, yet existing harnesses are hand-crafted by experts through slow build-and-test cycles. We propose VidHarness, a framework that automates harness design for cost-efficient long video understanding, in which a harness proposer iteratively evolves harnesses based on execution feedback from an evolution environment. To escape the local optima of greedy refinement, we organize the evolution as Monte Carlo tree search (MCTS), and to reduce the evaluation cost, we integrate uncertainty-aware multi-fidelity validation, which screens new harnesses on a few questions and promotes only the promising ones. Since the best harness varies with the frame budget, we further introduce a mixture-of-harness that routes each question to a harness specialized for its budget. VidHarness sets new state-of-the-art results on LongVideoBench, Video-MME, and Video-Holmes, outperforms the strongest hand-crafted video agent by up to 11.211.2 points, and generalizes to the knowledge-intensive benchmarks Video-MMMU and MMVU with fewer than half of the frames of uniform sampling.

Figures & tables

Explore similar work

CardsList
  1. MetaVideoAgent: Automated Video-Agent Evolution for Long-Form Video Understanding

    Aug 5, 2026Benlei Cui, Ruize Wang, Junjie Li +9Video AgentSelf-Evolving Agents

  2. Online Video Agent Harness for Long Video Understanding

    Sep 14, 2026Sen Yang, Boqiang Duan, Jing Yang +7Video AgentVideo Understanding