Text-to-Audio (T2A) retrievers are typically evaluated with caption style queries, but the same user intent can be expressed in many forms. We introduce CORA (Caption-Offset Retrieval for Audio), a caption anchored diagnostic protocol that rewrites each source caption into five intent preserving forms (Command, Question, Indirect, Key phrase, and Statement) while fixing the target audio. By tracking the same target across query forms, CORA defines RankDrop, a metric revealing failures hidden by Recall@k. Using Pearson's correlation coefficient r, we find that RankDrop is weakly associated with raw text space movement (r=0.084), but strongly associated with Target Alignment Loss and Target Boundary Margin Degradation (r=0.508 and r=0.615). The same pattern appears in OEA retrievers, where RankDrop is better explained by boundary degradation (r=0.472/0.478) than by query movement (r=0.084/0.046). Overall, these results suggest that robust T2A retrieval requires preserving the target's boundary advantage over competing audio under reformulation.
Figures & tables
Dataset
Source captions
Validation
Test
Excluded
Retained
Evaluations
AudioCaps
1,469
495
974
11
963
19,260
Clotho
2,090
1,045
1,045
51
994
19,880
MACS
1,000
500
500
0
500
10,000
MECAT
1,008
160
848
0
848
16,960
Total
5,567
2,200
3,367
62
3,305
66,100
Table 1: Dataset counts and filtering for the main CORA evaluation. Source captions include the validation and test splits, whereas retrieval uses only test items. The excluded items have an empty indirect query (11 AudioCaps items) or an empty key-phrase query (51 Clotho items). Each retained item contributes five query forms evaluated with four CLAP retrievers.
Figure 1: Overall MeanRD by query form, averaged across all datasets and models. Statement queries (most similar to original captions) are the most stable, while indirect queries cause the largest rank drop.
Dataset
Query form
MeanRD
AudioCaps
Command
-8.13
Indirect
11.63
Key phrase
23.58
Question
-13.83
Statement
-21.29
Clotho
Command
-1.07
Table 2: Dataset-by-form MeanRD. A query form that causes a large rank drop in one dataset can be neutral or beneficial in another, showing strong dependency on the dataset regime.
Dataset
Model
MeanRD
AudioCaps
LAION
49.63
M2D
22.79
MGA
31.10
MS-CLAP
-109.96
Clotho
LAION
31.92
M2D
17.00
Table 3: Dataset-by-model MeanRD under CORA reformulations. The direction and severity of rank drops vary across CLAP models within the same dataset, indicating that CORA sensitivity depends on the retriever as well as the dataset.
Figure 2: Correlation between diagnostic signals and RD over 66,100 CORA pairs using the pooled cross-dataset candidate set. Target-specific and boundary-specific signals, TAL and TBMD, show substantially stronger association with RD than query movement Δmove .
Selection policy
MeanRD
Average over views
3.881
Min- Δmove view
1.875
Min-TAL oracle
-18.568
Min-TBMD oracle
-19.467
Best-rank oracle
-24.560
Table 4: Controlled CORA-view intervention. Lower MeanRD indicates better target rank preservation; negative values indicate that the selected CORA view improves the target rank relative to the source query. Values in this table are used to compare selection policies under the same controlled setting.
Figure 3: Illustration of (a) the audio-projection and (b) the boundary-projection perspectives for query shift.
Feature set
Pearson r
Spearman ρ
fixed effects only
0.174
0.178
+ ∥d∥
0.206
0.196
+ paudio
0.500
0.594
+ pboundary
0.422
0.529
Table 5: Correlation between projection-based feature sets and RD prediction.
Model
CORA R@5
MeanRD
r(Δmove,RD)
r(TBMD,RD)
OEA-Qwen3B-Cl
43.34
0.603
0.084
0.472
OEA-Qwen3B-AC
41.17
0.479
0.046
0.478
RobustCLAP
35.47
9.408
0.174
0.508
Table 6: External retriever diagnostics under the CORA protocol. CORA R@5 is reported as a percentage. MeanRD measures average target-rank degradation after reformulation. Correlations are computed separately for each retriever.
Appendix figures & tables16 assets
Supplementary material from the paper’s appendix.
Appendix
Metric / Evaluation Axis
Value
Task 1 Raw Agreement (Semantic Alignment)
98.2%
Task 2 Accuracy (Query-Type Validity)
98.8%
Mean Pairwise Cohen’s κ
0.970
Pairwise Cohen’s κ Range
0.925–1.000
Appendix
Table 8: Summary of additional overlapping human validation across 5 annotators on the same 100 pairs.
Error Type
Confusion Pair
Count
Interrogative form
Question ↔ Indirect
3
Phrase vs. sentence
Key phrase ↔ Statement
1
Phrase vs. imperative
Key phrase ↔ Command
1
Declarative vs. indirect
Statement ↔ Indirect
1
Total
6
Appendix
Table 9: Breakdown of query-type confusion errors observed in the overlapping human validation.
Form
Real-user query (WildClaims)
CORA query
Key Phrase
Free resources to get AMHS based messages
Crowd speech laughter
Statement
The bomb was dropped on Hiroshima at around 8 am.
Many people are speaking and laughing.
Question
Can you find leet cheatsheet?
Does the recording feature pigeons cooing, bird wings flapping, gravel shuffling, and wood clacking?
Command
Show me sources behind all of these bullets.
Retrieve speech and laughter from many people.
Indirect
Can you please provide me with citable sources of journals that discusses the topic of organizational innovation technology adoption and the inconsistencies between research studies.
Would you be able to help me locate the audio where a train is running on train tracks and a steam engine horn is whistling?
Appendix
Table 10: Examples of real-user information-seeking queries from WildClaims and CORA queries exhibiting the same query forms.
Figure 4: Semantic diversity comparison between UIQ, CORA(w/o Prompt), and CORA. (a) Global semantic diversity across all query embeddings. (b) Within-type semantic diversity for command, question, and indirect queries. CORA consistently improves diversity over CORA(w/o Prompt), with the largest gain observed for the command type.
Query set
Forms
Evaluations
MeanRD
CORA
5
56,100
+5.28
CORA w/o div.
5
56,100
+4.17
Released UIQ
4
44,880
−5.78
Appendix
Table 11: MeanRD for three query sets evaluated on the same 2,805 items and per-dataset candidate pools.
Query set ( N )
Δmove
TAL
Dynamic-HN TBMD
CORA (56,100)
.092/.071
.457/.574
.535/.667
CORA w/o diversification (56,100)
.081/.064
.443/.578
.519/.660
Released UIQ (44,880)
−.075/−.073
.508/.669
.538/.736
Appendix
Table 12: Pearson r /Spearman ρ correlations with RankDrop under matched per-dataset candidate pools. N counts item-by-query-form-by-retriever observations across 2,805 items and four retrievers. This is a descriptive comparison because query generation, grounding, and query-type inventories differ.
Retriever
Type
Source R@5
CORA R@5
Δ R@5
MeanRD
LAION
single
29.68
22.17
-7.51
104.71
MS-CLAP
single
17.37
20.44
3.07
-2.67
M2D
single
34.89
29.57
-5.31
2.46
MGA
single
34.52
31.69
-2.83
30.34
MultiCLAP-ScoreAvg-Z
fusion
38.67
35.75
-2.92
24.29
MultiCLAP-ScoreAvg-MinMax
fusion
38.28
35.50
-2.77
25.04
Appendix
Table 13: Comprehensive retriever-level performance under the CORA protocol. R@5 and Δ R@5 values are reported as percentages. MeanRD is computed over all CORA pairs. This table illustrates the limitation of aggregate Δ R@5 and motivates the use of paired MeanRD to accurately capture instance-level rank degradation.
Feature terms
Pearson r
Spearman ρ
additive model/dataset/query
0.240
0.244
model-dataset + query
0.285
0.257
pairwise query interactions
0.309
0.317
full model-dataset-query cell
0.376
0.391
Appendix
Table 18: Grouped cross-validated prediction of RD. Interaction terms improve prediction over additive model, dataset, and query form terms, indicating that CORA sensitivity is conditional on the model-dataset regime.
Top-1 boundary degradation definition
Pearson w/ RD
CORA-HN fixed TBMD
0.487
Source-HN fixed TBMD
0.317
Dynamic top-1 TBMD (pooled)
0.598
Appendix
Table 19: Sensitivity of TBMD to top-1 hard negative anchoring choices. All variants use top-1 non-target audios, matching the main diagnostic setting. The dynamic top-1 definition used in the main text gives the strongest association with RD.
Anchor/control
N
r(TAL,RD)
Source query → CORA query
66,100
0.615
Statement query → CORA4
45,888
0.399
Content-fixed wrapper control
32,320
0.406
Appendix
Table 20: Anchor and content controls for Target Alignment Loss (TAL), computed with the pooled cross-dataset candidate set. TAL is degradation-oriented, so positive correlations indicate that larger loss of target alignment is associated with larger RD.
Analysis
Per-dataset
Pooled
TAL-only (Pearson r )
0.561
0.615
TAL + dynamic-HN TBMD
0.644
0.626
Increment
+0.083
+0.010
Partial r beyond TAL
0.382
0.144
Appendix
Table 21: Control analysis for TBMD. Adding dynamic-HN TBMD to TAL improves RD prediction in both settings, and TBMD keeps a positive partial correlation with RD once TAL, dataset, model, query form, Δmove , source rank, and pool size are held fixed.
Model
RankDrop group
N
MeanRD
Content+
Content-
Form-
Intent-
Mean margin
LAION
worsened
160
71.819
0.2832
0.2387
0.1205
0.0521
-0.2189
LAION
improved/stable
160
-37.863
0.2799
0.2647
0.1143
0.0409
-0.0928
MS-CLAP
worsened
160
45.581
0.2774
0.2840
0.1025
0.0486
-0.1250
MS-CLAP
improved/stable
160
-50.213
0.2696
0.2885
0.1191
0.0389
-0.0697
M2D
worsened
160
31.356
0.2780
0.2915
0.1178
0.0404
-0.0430
M2D
improved/stable
160
-23.938
0.3025
0.2710
0.1243
0.0466
-0.0169
Appendix
Table 22: Token-group saliency shares by RD group. Mean margin is the local target boundary margin under the CORA query. Across all four models, rank-worsened pairs have lower mean boundary margins than rank-improved or stable pairs.
Feature set
Pearson r
Spearman ρ
Fixed effects only
0.077
0.007
Query movement
0.115
0.045
Token-group saliency
0.090
0.037
Boundary features
0.634
0.591
Boundary + token saliency
0.632
0.581
Appendix
Table 23: Grouped cross-validated RD prediction from token saliency and boundary features. Token-group saliency alone has weak predictive value, and adding it to boundary features does not improve prediction.
Dataset
N
MeanRD
Worsened (%)
r(Δmove)
r(TAL)
r(TBMD)
AudioCaps
19,260
-1.61
43.2
-0.016
0.531
0.547
Clotho
19,880
27.71
51.6
0.221
0.580
0.585
MACS
10,000
49.69
50.2
0.361
0.737
0.716
MECAT
16,960
71.43
55.5
0.216
0.700
0.668
Appendix
Table 24: Dataset-level diagnostic correlations with RD. Worsened denotes the percentage of pairs for which the CORA-query rank is lower than the source-query rank. TAL and TBMD are consistently more associated with RD than query movement across datasets.
Retriever
Source R@5
CORA R@5
Δ R@5
MeanRD
M2D
34.89
29.57
-5.31
2.46
MGA
34.52
31.69
-2.83
30.34
MultiCLAP-ScoreAvg-Z
38.67
35.75
-2.92
24.29
MultiCLAP-ScoreAvg-MinMax
38.28
35.50
-2.77
25.04
Appendix
Table 25: Additional inference-time score fusion check under the CORA protocol. M2D and MGA are included as strong single-retriever references. R@5 and Δ R@5 values are reported as percentages.
Condition
RD
R@5
Δ move
TAL
TBMD
No canonicalizer
6.33
37.33
0.103
0.551
0.653
General rewriting
4.14
38.10
0.080
0.519
0.634
CORA-aware
-0.03
40.02
0.020
0.546
0.683
Appendix
Table 26: Effect of query preprocessing on retrieval performance and diagnostic associations. RD denotes Mean RankDrop, and R@5 is reported in percent. Δ move, TAL, and TBMD report Pearson’s r with RD. Source-query R@5 is 38.95% in all conditions.
School of Software, Nanjing University Nanjing, China · School of Software, Northwestern Polytechnical University Xi’an, China · College of Computing and Data Science, Nanyang Technological University Singapore, Singapore