cs.CVSep 27, 2026

OPERA: A Unified Omnimodal Progressive Spatio-Temporal Reasoning Agent for Referring Video Segmentation

Authors: Jingchen Ni, Yuji Wang, Shannan Yan, Haoru Li, Sitong Chen, Chun Yuan

Organizations: Tsinghua University · University of California San Diego · ETH Zurich

Abstract

Referring video segmentation with heterogeneous multimodal queries---spanning text, audio, and reference images---demands both robust cross-modal understanding and precise spatio-temporal reasoning. We propose OPERA (Omnimodal Progressive spatio-tEmporal Reasoning Agent), a unified reasoning agent built on a single MLLM that performs dual-axis progressive reasoning via three specialized stages. Along the temporal axis, a Temporal Reasoning Agent narrows the frame search space through coarse-to-fine filtering to identify the most informative key frame. Along the spatial axis, a Distillation Agent establishes what to locate via cross-modal semantic distillation, and a Grounding Agent enhanced with GRPO determines where the target appears, with dense mask propagation completing the pixel-level output. OPERA sets a new state of the art on OmniAVS and Ref-AVS and transfers zero-shot to standard referring video segmentation benchmarks.

Figures & tables

Appendix figures & tables1 asset

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. ReflexTrack: A Feedback-Driven Agent for Training-Free Referring Video Object Segmentation

    Jul 27, 2026Yuanjia Li, Tianyang Xu, Tao Zhou +3Video Object SegmentationSpatial Grounding

  2. SteerSeg: Attention Steering for Reasoning Video Segmentation

    May 14, 2026Ali Cheraghian, Hamidreza Dastmalchi, Abdelwahed Khamis +3Video Object SegmentationVision-Language Model Grounding

  3. Event-Aware Instructed Assistant for Referring Video Segmentation

    Jun 25, 2026Jinyu Liu, Henghui Ding, Shuting He +1Video Object SegmentationHierarchical