CORA: A Protocol for Diagnosing Boundary Robustness in Text-to-Audio Retrieval under Query Reformulations
Organizations: Sogang University
Abstract
Text-to-Audio (T2A) retrievers are typically evaluated with caption style queries, but the same user intent can be expressed in many forms. We introduce CORA (Caption-Offset Retrieval for Audio), a caption anchored diagnostic protocol that rewrites each source caption into five intent preserving forms (Command, Question, Indirect, Key phrase, and Statement) while fixing the target audio. By tracking the same target across query forms, CORA defines RankDrop, a metric revealing failures hidden by Recall@k. Using Pearson's correlation coefficient r, we find that RankDrop is weakly associated with raw text space movement (r=0.084), but strongly associated with Target Alignment Loss and Target Boundary Margin Degradation (r=0.508 and r=0.615). The same pattern appears in OEA retrievers, where RankDrop is better explained by boundary degradation (r=0.472/0.478) than by query movement (r=0.084/0.046). Overall, these results suggest that robust T2A retrieval requires preserving the target's boundary advantage over competing audio under reformulation.
Figures & tables
| Dataset | Source captions | Validation | Test | Excluded | Retained | Evaluations |
|---|---|---|---|---|---|---|
| AudioCaps | 1,469 | 495 | 974 | 11 | 963 | 19,260 |
| Clotho | 2,090 | 1,045 | 1,045 | 51 | 994 | 19,880 |
| MACS | 1,000 | 500 | 500 | 0 | 500 | 10,000 |
| MECAT | 1,008 | 160 | 848 | 0 | 848 | 16,960 |
| Total | 5,567 | 2,200 | 3,367 | 62 | 3,305 | 66,100 |
| Dataset | Query form | MeanRD |
|---|---|---|
| AudioCaps | Command | -8.13 |
| Indirect | 11.63 | |
| Key phrase | 23.58 | |
| Question | -13.83 | |
| Statement | -21.29 | |
| Clotho | Command | -1.07 |
| Dataset | Model | MeanRD |
|---|---|---|
| AudioCaps | LAION | 49.63 |
| M2D | 22.79 | |
| MGA | 31.10 | |
| MS-CLAP | -109.96 | |
| Clotho | LAION | 31.92 |
| M2D | 17.00 |
| Selection policy | MeanRD |
|---|---|
| Average over views | 3.881 |
| Min- view | 1.875 |
| Min-TAL oracle | -18.568 |
| Min-TBMD oracle | -19.467 |
| Best-rank oracle | -24.560 |
| Feature set | Pearson | Spearman |
|---|---|---|
| fixed effects only | 0.174 | 0.178 |
| + | 0.206 | 0.196 |
| + | 0.500 | 0.594 |
| + | 0.422 | 0.529 |
| Model | CORA R@5 | MeanRD | ||
|---|---|---|---|---|
| OEA-Qwen3B-Cl | 43.34 | 0.603 | 0.084 | 0.472 |
| OEA-Qwen3B-AC | 41.17 | 0.479 | 0.046 | 0.478 |
| RobustCLAP | 35.47 | 9.408 | 0.174 | 0.508 |
Appendix figures & tables16 assets
Supplementary material from the paper’s appendix.
Appendix
| Metric / Evaluation Axis | Value |
|---|---|
| Task 1 Raw Agreement (Semantic Alignment) | 98.2% |
| Task 2 Accuracy (Query-Type Validity) | 98.8% |
| Mean Pairwise Cohen’s | 0.970 |
| Pairwise Cohen’s Range | 0.925–1.000 |
| Error Type | Confusion Pair | Count |
|---|---|---|
| Interrogative form | Question Indirect | 3 |
| Phrase vs. sentence | Key phrase Statement | 1 |
| Phrase vs. imperative | Key phrase Command | 1 |
| Declarative vs. indirect | Statement Indirect | 1 |
| Total | 6 |
| Form | Real-user query (WildClaims) | CORA query |
|---|---|---|
| Key Phrase | Free resources to get AMHS based messages | Crowd speech laughter |
| Statement | The bomb was dropped on Hiroshima at around 8 am. | Many people are speaking and laughing. |
| Question | Can you find leet cheatsheet? | Does the recording feature pigeons cooing, bird wings flapping, gravel shuffling, and wood clacking? |
| Command | Show me sources behind all of these bullets. | Retrieve speech and laughter from many people. |
| Indirect | Can you please provide me with citable sources of journals that discusses the topic of organizational innovation technology adoption and the inconsistencies between research studies. | Would you be able to help me locate the audio where a train is running on train tracks and a steam engine horn is whistling? |
| Query set | Forms | Evaluations | MeanRD |
|---|---|---|---|
| CORA | 5 | 56,100 | +5.28 |
| CORA w/o div. | 5 | 56,100 | +4.17 |
| Released UIQ | 4 | 44,880 |
| Query set ( ) | TAL | Dynamic-HN TBMD | |
|---|---|---|---|
| CORA (56,100) | .092/.071 | .457/.574 | .535/.667 |
| CORA w/o diversification (56,100) | .081/.064 | .443/.578 | .519/.660 |
| Released UIQ (44,880) | .508/.669 | .538/.736 |
| Retriever | Type | Source R@5 | CORA R@5 | R@5 | MeanRD |
|---|---|---|---|---|---|
| LAION | single | 29.68 | 22.17 | -7.51 | 104.71 |
| MS-CLAP | single | 17.37 | 20.44 | 3.07 | -2.67 |
| M2D | single | 34.89 | 29.57 | -5.31 | 2.46 |
| MGA | single | 34.52 | 31.69 | -2.83 | 30.34 |
| MultiCLAP-ScoreAvg-Z | fusion | 38.67 | 35.75 | -2.92 | 24.29 |
| MultiCLAP-ScoreAvg-MinMax | fusion | 38.28 | 35.50 | -2.77 | 25.04 |
| Feature terms | Pearson | Spearman |
|---|---|---|
| additive model/dataset/query | 0.240 | 0.244 |
| model-dataset + query | 0.285 | 0.257 |
| pairwise query interactions | 0.309 | 0.317 |
| full model-dataset-query cell | 0.376 | 0.391 |
| Top-1 boundary degradation definition | Pearson w/ RD |
|---|---|
| CORA-HN fixed TBMD | 0.487 |
| Source-HN fixed TBMD | 0.317 |
| Dynamic top-1 TBMD (pooled) | 0.598 |
| Anchor/control | N | |
|---|---|---|
| Source query CORA query | 66,100 | 0.615 |
| Statement query CORA4 | 45,888 | 0.399 |
| Content-fixed wrapper control | 32,320 | 0.406 |
| Analysis | Per-dataset | Pooled |
|---|---|---|
| TAL-only (Pearson ) | 0.561 | 0.615 |
| TAL + dynamic-HN TBMD | 0.644 | 0.626 |
| Increment | +0.083 | +0.010 |
| Partial beyond TAL | 0.382 | 0.144 |
| Model | RankDrop group | N | MeanRD | Content+ | Content- | Form- | Intent- | Mean margin |
|---|---|---|---|---|---|---|---|---|
| LAION | worsened | 160 | 71.819 | 0.2832 | 0.2387 | 0.1205 | 0.0521 | -0.2189 |
| LAION | improved/stable | 160 | -37.863 | 0.2799 | 0.2647 | 0.1143 | 0.0409 | -0.0928 |
| MS-CLAP | worsened | 160 | 45.581 | 0.2774 | 0.2840 | 0.1025 | 0.0486 | -0.1250 |
| MS-CLAP | improved/stable | 160 | -50.213 | 0.2696 | 0.2885 | 0.1191 | 0.0389 | -0.0697 |
| M2D | worsened | 160 | 31.356 | 0.2780 | 0.2915 | 0.1178 | 0.0404 | -0.0430 |
| M2D | improved/stable | 160 | -23.938 | 0.3025 | 0.2710 | 0.1243 | 0.0466 | -0.0169 |
| Feature set | Pearson | Spearman |
|---|---|---|
| Fixed effects only | 0.077 | 0.007 |
| Query movement | 0.115 | 0.045 |
| Token-group saliency | 0.090 | 0.037 |
| Boundary features | 0.634 | 0.591 |
| Boundary + token saliency | 0.632 | 0.581 |
| Dataset | N | MeanRD | Worsened (%) | |||
|---|---|---|---|---|---|---|
| AudioCaps | 19,260 | -1.61 | 43.2 | -0.016 | 0.531 | 0.547 |
| Clotho | 19,880 | 27.71 | 51.6 | 0.221 | 0.580 | 0.585 |
| MACS | 10,000 | 49.69 | 50.2 | 0.361 | 0.737 | 0.716 |
| MECAT | 16,960 | 71.43 | 55.5 | 0.216 | 0.700 | 0.668 |
| Retriever | Source R@5 | CORA R@5 | R@5 | MeanRD |
|---|---|---|---|---|
| M2D | 34.89 | 29.57 | -5.31 | 2.46 |
| MGA | 34.52 | 31.69 | -2.83 | 30.34 |
| MultiCLAP-ScoreAvg-Z | 38.67 | 35.75 | -2.92 | 24.29 |
| MultiCLAP-ScoreAvg-MinMax | 38.28 | 35.50 | -2.77 | 25.04 |
| Condition | RD | R@5 | move | TAL | TBMD |
|---|---|---|---|---|---|
| No canonicalizer | 6.33 | 37.33 | 0.103 | 0.551 | 0.653 |
| General rewriting | 4.14 | 38.10 | 0.080 | 0.519 | 0.634 |
| CORA-aware | -0.03 | 40.02 | 0.020 | 0.546 | 0.683 |