cs.CVMay 2, 2026

Two-Pass Zero-Shot Temporal-Spatial Grounding of Rare Traffic Events in Surveillance Video

Authors: Jiantang Huang

Abstract

Grounding traffic accidents in real CCTV footage is a rare-event problem where training on labeled accident video is often prohibited, yet accurate joint localization in time, space, and collision type is required. We present a no-fine-tuning pipeline that elicits this joint output from frozen vision-language models through two ideas. First, a coarse-to-fine two-pass decomposition: a full-video pass at 1 fps produces a coarse (t, x, y, c) tuple, then a second pass at 5 fps within a +/- 3 s window refines time and location, with two deterministic confidence gates that revert to the coarse estimate on boundary hedges or edge-clamped coordinates. Second, a specialist role assignment: Qwen3-VL-Plus handles grounding, Gemini 3.1 Flash-Lite handles typing on a centered video clip. On the ACCIDENT@CVPR 2026 benchmark (2,027 real CCTV videos) we reach ACC^S = 0.539 (95% CI [0.525, 0.553]): +0.127 over the benchmark paper's best-of-baselines oracle (0.412), +0.143 over the strongest single-VLM baseline (Molmo-7B, 0.396), and +0.250 over the naive baseline (0.289). The VLM path uses up to three API calls per video (17% fall back to physics on API failures); the full run costs ~$20.

Explore similar work

CardsList
  1. Zero-Shot Traffic Accident Detection via a Coarse-to-Fine VLM-Tracking Pipeline

    Aug 9, 2026Dipit Saha, Shah Mohammad Abdul Mannan, Mohammad Raihan Rashid +2AccidentsRobust Vehicle Recognition

  2. Metadata-Aware Multi-Prompt Reasoning for Zero-Shot Accident Understanding

    Jun 10, 2026Tarandeep Singh, Soumyanetra Pal, Soham Biswas +1Accidents