CueRator: Agentic Search for Symbolic Rules to Adapt Frozen Multimodal Encoders
Organizations: Pusan National University
Abstract
Large language model agents have been used to search over symbolic structures such as programs and equations. We propose CueRator, an agentic framework for policy-aware decision-rule discovery, which adapts frozen contrastive multimodal encoders by searching for the decision rule that converts their cross-modal similarities into predictions. We validate it on open-vocabulary audio-visual event perception, where existing methods involve a trade-off between adaptivity and generalization to unseen categories: trained modules adapt at the cost of generalization, and fixed rules the reverse. The framework pairs a symbolic formulation for generalization with a lightweight policy that predicts its parameters per video for adaptivity. A report-guided multi-agent loop discovers the formulation offline, evaluating each candidate on its expressive ceiling and on whether a trained policy can realize it. On OV-AVEBench, CueRator raises the total average from 57.8 to 60.2 and unseen-category performance from 55.8 to 59.9 over the best existing method, reducing the seen-unseen gap from 7.1 to 1.2. Ablations attribute the gains to both the formulation and the policy and show that both feedback signals are necessary for effective search. CueRator also improves over the respective baselines on two further audio-visual event perception tasks, and the discovered rule remains competitive across encoders with only the policy retrained. Code is available at https://github.com/cvsp-lab/cuerator.
Figures & tables
| Seen | Unseen | Total | ||||||||||
| Method | Acc. | Seg. | Eve. | Avg. | Acc. | Seg. | Eve. | Avg. | Acc. | Seg. | Eve. | Avg. |
| Training-free | ||||||||||||
| Video-LLaMA2 | 50.1 | 40.6 | 32.0 | 40.9 | 48.5 | 38.5 | 29.0 | 38.6 | 48.9 | 39.1 | 29.8 | 39.3 |
| CLIP&CLAP | 51.4 | 41.4 | 31.9 | 41.6 | 51.6 | 42.2 | 31.6 | 41.8 | 51.5 | 41.9 | 31.7 | 41.7 |
| ImageBind-TF | 57.5 | 45.0 | 34.0 | 45.5 | 59.8 | 47.3 | 34.0 | 47.0 | 59.2 | 46.7 | 34.0 | 46.6 |
| Fine-tuning | ||||||||||||
| Variant | Oracle | Test | |
|---|---|---|---|
| Acc. | Acc. | Avg. | |
| CueRator (full) | 88.7 0.9 | 68.3 0.8 | 59.4 0.9 |
| w/o policy agent | 91.5 2.3 | 66.7 0.9 | 56.3 1.7 |
| w/o oracle agent | 87.3 5.0 | 66.0 0.2 | 55.3 0.5 |
| score-only loop | 88.4 1.3 | 64.1 1.1 | 53.6 0.7 |
| no feedback | 86.0 3.5 | 62.2 2.2 | 51.8 2.2 |
| Benchmark | Baseline | Metric | Seen:Unseen | Performance | |
|---|---|---|---|---|---|
| Baseline | CueRator | ||||
| OV-DAVEL | Open-DAVTR | mAP | 75:25 | 22.9 | 32.0 |
| 50:50 | 19.4 | 32.7 | |||
| OV-AVVP | AV 2 A | Seg. Type@AV | 17:8 | 52.4 | 55.6 |
Appendix figures & tables26 assets
Supplementary material from the paper’s appendix.
Appendix
| Component | Time (h) | Share |
|---|---|---|
| Agent (LLM inference) | 4.0 | 41% |
| Oracle search | 1.8 | 18% |
| Policy training & evaluation | 4.0 | 41% |
| Total | 9.8 | 100% |
| Term | Meaning |
|---|---|
| Formulation | A closed-form rule computing per-segment thresholds from the audio, visual, and text embeddings (via their similarities), with tunable parameters . |
| Oracle ceiling | The performance of when is set optimally for each video—an upper bound on what the rule’s structure can express. |
| Policy network | A small network predicting from the audio and visual embeddings, called a policy because it is trained with REINFORCE (thresholding is non-differentiable). |
| Learnability | Whether the optimal is predictable from the video—measured as the gap between the oracle ceiling and the policy’s performance. |
| Reports / Memory / Directives | The agents’ written analyses ( reports ) accumulate across iterations ( memory ); the Plan Agent reads them and issues instructions for the next iteration ( directives ). |
| Agent | Receives | Performs | Writes to memory |
|---|---|---|---|
| Plan | Prior candidate summaries, per-candidate records (formulations, parameter ranges, oracle scores, policy validation scores, agent reports). | Diagnoses the search state; decides whether to refine, abandon, or redirect formulation families; assigns parameter roles, range changes, ablation targets, and analysis requests. | Design directive, oracle ablation protocol, policy analysis protocol, and one-line planning summary. |
| Design | Plan directive, formulation interface, symbolic-form constraints, and prior formulation code patterns. | Proposes an executable formulation ; declares parameter roles and ranges ; revises the code until AST checks pass. | Candidate code, parameter ranges, parameter role descriptions, validity status, and design summary. |
| Oracle | Verified formulation , parameter ranges , training split, and requested ablations. | Searches per-video oracle parameters; estimates expressive capacity; repeats per-video oracle search for each ablation; analyzes parameter distributions, boundary saturation, and component impact. | Oracle score , ablation statistics, parameter diagnostics, raw-analysis artifacts, and oracle report. |
| Policy | Verified formulation , parameter ranges , trained policy outputs, validation metrics, and requested diagnostics. | Analyzes the policy trained for the candidate: policy-instantiated behavior, predicted parameter distributions, and failure modes. | Validation score , predicted-parameter statistics, learnability diagnostics, and policy report. |
| Parameter | Multiplied term | Range | Code index |
|---|---|---|---|
| audio bias | params[0] | ||
| params[2] | |||
| params[1] | |||
| params[3] | |||
| params[4] | |||
| visual bias | params[5] |
| Cue group | Terms | Interpretation |
|---|---|---|
| Cross-modal support | , | Each modality’s threshold uses the other modality’s category ranking at the same segment. |
| Local AV agreement | , | favors categories supported by both modalities. |
| Clip-level competition | , | Compares the category against its competitors over the full video. |
| Temporal calibration | , | Compares the current segment with the clip-level evidence for the same category. |
| Encoder | AV 2 A | Transferred |
|---|---|---|
| LanguageBind | 50.6 | 58.6 |
| CLIP+CLAP | 48.9 | 54.5 |
| Benchmark | Metric | Baseline | Task-specific | Transferred |
|---|---|---|---|---|
| OV-DAVEL (50:50) | mAP | 19.4 | 32.7 | 30.3 |
| OV-AVVP (17:8) | Seg. Type@AV | 52.4 | 55.6 | 49.5 |
| Backbone | Access | Test Avg. |
|---|---|---|
| GPT-5.6 Sol | Closed | 59.5 0.9 |
| GPT-5.4 | Closed | 59.4 0.9 |
| Claude Opus 4.8 | Closed | 57.4 1.1 |
| Qwen3.7-Max | Closed | 57.1 1.8 |
| GLM-5.2 | Open-weight | 56.7 1.0 |
| Setting | Oracle Acc. | Val Acc. | Test Avg. |
|---|---|---|---|
| CueRator (default) | 88.7 0.9 | 67.8 0.2 | 59.4 0.9 |
| w/o linear form | 88.4 0.3 | 65.2 0.6 | 55.1 0.8 |
| w/o constant set | 85.6 4.1 | 66.4 0.8 | 56.6 1.7 |
| w/o statement budget | 86.6 5.6 | 66.6 0.5 | 57.0 0.2 |
| Terms per modality (incl. bias) | Search time (h/session) | Test Avg. | |
|---|---|---|---|
| 6 | 3 | 8.2 0.2 | 54.6 0.3 |
| 8 | 4 | 9.0 0.2 | 58.3 1.1 |
| 10 | 5 | 9.8 0.2 | 59.4 0.9 |
| 12 | 6 | 10.0 0.1 | 57.2 1.0 |
| 14 | 7 | 10.6 0.4 | 55.3 1.9 |
| Reward | Acc. | Seg. | Eve. | Avg. |
|---|---|---|---|---|
| Acc. (default) | 68.6 | 59.4 | 52.7 | 60.2 |
| Metric average | 67.2 | 59.2 | 53.0 | 59.8 |