cs.CVOct 6, 2026

CueRator: Agentic Search for Symbolic Rules to Adapt Frozen Multimodal Encoders

Authors: Sunchan Park, Beomkwon Cho, Kyeongbo Kong

Organizations: Pusan National University

Abstract

Large language model agents have been used to search over symbolic structures such as programs and equations. We propose CueRator, an agentic framework for policy-aware decision-rule discovery, which adapts frozen contrastive multimodal encoders by searching for the decision rule that converts their cross-modal similarities into predictions. We validate it on open-vocabulary audio-visual event perception, where existing methods involve a trade-off between adaptivity and generalization to unseen categories: trained modules adapt at the cost of generalization, and fixed rules the reverse. The framework pairs a symbolic formulation for generalization with a lightweight policy that predicts its parameters per video for adaptivity. A report-guided multi-agent loop discovers the formulation offline, evaluating each candidate on its expressive ceiling and on whether a trained policy can realize it. On OV-AVEBench, CueRator raises the total average from 57.8 to 60.2 and unseen-category performance from 55.8 to 59.9 over the best existing method, reducing the seen-unseen gap from 7.1 to 1.2. Ablations attribute the gains to both the formulation and the policy and show that both feedback signals are necessary for effective search. CueRator also improves over the respective baselines on two further audio-visual event perception tasks, and the discovered rule remains competitive across encoders with only the policy retrained. Code is available at https://github.com/cvsp-lab/cuerator.

Figures & tables

Appendix figures & tables26 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. OmniSeek: Native Tool Integration for Multi-turn Audio-Visual Reasoning

    Oct 1, 2026Haibo Wang, Jiteng Mu, Jialu Li +5Audio-Visual ReasoningOmni-Modal

  2. AgentCVR: Active Multi-Agent Cross-Video Reasoning via Script-Simulated Reinforcement Learning

    May 28, 2026Yilun Qiu, Jiahe Wang, Cilin Yan +4Multimodal AgentsSimulation-Based Reinforcement Learning

  3. A Strong Baseline for Evaluating Vision Encoders in Multimodal Large Language Models

    Oct 4, 2026Yilin Yang, Jun-Tao Tang, Kengyi Wang +3Vision Encoders