Vision-language models (VLMs) can answer questions about hour-long videos, but processing every frame is prohibitively expensive, even though the evidence for a question usually spans only a few seconds. Video agents, i.e., harness programs wrapped around a frozen VLM, address this by observing the video selectively, yet existing harnesses are hand-crafted by experts through slow build-and-test cycles. We propose VidHarness, a framework that automates harness design for cost-efficient long video understanding, in which a harness proposer iteratively evolves harnesses based on execution feedback from an evolution environment. To escape the local optima of greedy refinement, we organize the evolution as Monte Carlo tree search (MCTS), and to reduce the evaluation cost, we integrate uncertainty-aware multi-fidelity validation, which screens new harnesses on a few questions and promotes only the promising ones. Since the best harness varies with the frame budget, we further introduce a mixture-of-harness that routes each question to a harness specialized for its budget. VidHarness sets new state-of-the-art results on LongVideoBench, Video-MME, and Video-Holmes, outperforms the strongest hand-crafted video agent by up to 11.2 points, and generalizes to the knowledge-intensive benchmarks Video-MMMU and MMVU with fewer than half of the frames of uniform sampling.
Figures & tables
Figure 1: Overview of harness evolution in VidHarness. Left: in the evolution environment, the harness proposer follows the evolution guidance and uses the interaction APIs to design new harnesses, whose evolution feedback is returned to the proposer for the next iteration. Right: comparison of evolution algorithms, where each circle is a harness and its bar shows the fraction of validation questions evaluated. Vanilla iterative refinement grows a single lineage and MCTS grows a tree, but both validate every harness in full; our multi-fidelity MCTS validates new harnesses on a small subset and re-validates only the promising ones.
Figure 2: Validation set statistics and mixture-of-harness. Left: distributions of data sources, video durations, transcription availability, and question types (multiple-choice, MCQ, and open-ended, OE) in the validation set. Right: a single top-ranked harness serves questions of all frame budgets but fits only some of them well (top), whereas the mixture-of-harness routes each question to the harness specialized for its budget tier (bottom).
Table 3: Ablation studies on Video-Holmes. Top: evolution strategies, where Evolution Time is the total evolution time and Val. Samples is the total number of validation samples consumed during evolution. Bottom: mixture-of-harness (MoH) versus the top-1 harness from evolution at each of the four budget tiers (low, medium, high, and xhigh).