Composed image retrieval (CIR) aims to retrieve a desired target image from a query consisting of a reference image and a modification text. This task exhibits an unusual representational asymmetry: the modification text specifies a transition from the reference state, whereas retrieval candidates depict completed target states. This creates a representation mismatch for zero-shot methods that query pretrained vision-language spaces directly with transformation-oriented language. We study this mismatch and reformulate zero-shot composed image retrieval as target-state reconstruction followed by retrieval. We instantiate this formulation with ASAP-CIR, a training-free framework that reconstructs a static target representation using a frozen multimodal large language model (MLLM). The representation combines multiple holistic descriptions with a variable set of importance-weighted atomic semantics, thereby preserving both overall target identity and fine-grained visual constraints. Retrieval then integrates holistic state alignment, atomic constraint grounding, and calibrated target-state evidence aggregation. A controlled text-only diagnostic shows that target-side static query formulations achieve more reliable retrieval than dynamic composed query formulations, particularly when source-state semantics must be suppressed or transformed. Experiments on FashionIQ, CIRR, and CIRCO further characterize the effectiveness and limitations of this representation principle, with the clearest gains on the multi-target CIRCO benchmark. These results show that how composed intent is represented before retrieval is a consequential design choice, distinct from the choice of retrieval backbone itself.
Figures & tables
Figure 1: Motivation of ASAP-CIR. The left panels compare three query formulation paradigms: (a) pseudo-word inversion directly combines an inverted visual token with dynamic modification text; (b) MLLM-based reasoning generates a target description without explicitly constraining it to static visual semantics; and (c) ASAP-CIR reconstructs the composed query as complementary holistic descriptions and importance-weighted atomic semantics for multi-level retrieval. Panel (d) visualizes the phenomenon observed in our diagnostic experiments: target-side static query embeddings are more closely aligned with the target embedding region, whereas dynamic composed query embeddings are more dispersed and frequently deviate from the intended target state.
Figure 2: Overview of the proposed ASAP-CIR framework. Given a composed query, the MLLM parser performs target-state semantic reconstruction and produces complementary holistic descriptions and importance-weighted atomic semantics. In the figure, these outputs are labeled “Holistic Descriptions” and “Weighted Atomic Semantics,” respectively. The holistic descriptions are encoded independently and combined by average pooling before being aligned with global image features to obtain the holistic evidence score. The atomic semantics are encoded independently and grounded against candidate patch features through patch-wise softmax assignment and importance-weighted aggregation, producing the atomic evidence score. Finally, query-wise z-score normalization calibrates both evidence scores, and weighted linear fusion produces the final retrieval score and ranking.
Backbone
Method
Venue
T-Free
Dress
Shirt
Toptee
Average
R@10
R@50
R@10
R@50
R@10
R@50
R@10
R@50
ViT-B/32
SEARLE
ICCV’23
N
18.54
39.51
24.44
41.61
25.70
46.46
22.89
42.53
FlowCIR
ECCV’26
N
24.80
41.80
19.10
40.10
26.40
47.80
23.40
43.20
CIReVL
ICLR’24
Y
25.29
46.36
28.36
47.84
31.21
53.85
28.29
49.35
LDRE
SIGIR’24
Y
19.97
41.84
27.38
46.27
27.07
48.78
24.81
45.63
AutoCIR
KDD’25
Y
26.52
46.36
32.43
51.67
33.96
56.09
30.97
51.37
Table 1: FashionIQ validation results. The best and second-best results within each backbone group are shown in bold and underlined, respectively.
Backbone
Method
Venue
T-Free
R@1
R@5
R@10
R@50
ViT-B/32
SEARLE
ICCV’23
N
24.00
53.42
66.82
89.78
TT-RLDR
AAAI’26
N
27.91
–
71.98
91.89
FlowCIR
ECCV’26
N
25.50
56.50
69.80
–
CIReVL
ICLR’24
Y
23.94
52.51
66.00
86.95
LDRE
SIGIR’24
Y
25.69
55.13
69.04
89.90
OSrCIR
CVPR’25
Y
25.42
54.54
68.19
–
Table 2: CIRR test1 results. The best and second-best results within each backbone group are shown in bold and underlined, respectively.
Backbone
Method
Venue
T-Free
mAP@5
mAP@10
mAP@25
mAP@50
ViT-L/14
Pic2Word
CVPR’23
N
8.72
9.51
10.64
11.29
SEARLE
ICCV’23
N
11.68
12.73
14.33
15.12
LinCIR
CVPR’24
N
12.62
13.40
14.81
15.69
DiffComp
CVPR’26
N
16.19
17.32
19.24
20.30
FlowCIR
ECCV’26
N
14.90
15.70
17.30
18.20
CIReVL
ICLR’24
Y
18.57
19.01
20.89
21.80
Table 3: CIRCO test results. The best and second-best results within each backbone group are shown in bold and underlined, respectively.
Backbone
Static R@1
Dynamic R@1
Gap
ViT-B/32
81.52
29.69
51.83
ViT-L/14
83.23
29.00
54.22
ViT-G/14
86.21
32.59
53.61
Table 4: Controlled text-only retrieval results for target-side static and dynamic composed query formulations under different OpenCLIP backbones. Gap denotes the R@1 difference between the two formulations.
Figure 3: R@1 comparison between target-side static and dynamic composed query formulations under the default OpenCLIP ViT-L/14 text encoder. The synthetic evaluation set contains 4,096 Qwen3-32B generated samples across eight dynamic semantic categories.
Figure 4: Target-anchored principal component analysis (PCA) projection of CLIP text features under the default OpenCLIP ViT-L/14 text encoder. For each of the eight operation types, 50 query samples are randomly selected. Squares denote target-side static query features, crosses denote dynamic composed query features, and triangles denote target features.
Matching evidence
Importance weighting
mAP@5
mAP@10
mAP@25
mAP@50
Holistic only
–
35.19
36.14
38.88
40.03
Atomic only
×
11.96
12.61
13.92
14.66
Atomic only
✓
13.19
13.84
15.26
16.07
Holistic + atomic
×
35.40
36.39
39.10
40.25
Holistic + atomic
✓
35.49
36.46
39.21
40.35
Table 5: Component ablation of holistic state alignment, atomic constraint grounding, and importance weighting on CIRCO with ViT-L/14.
Variant
MLLM parser
mAP@5
mAP@10
mAP@25
mAP@50
ASAP-CIR
Qwen3.6-27B
35.49
36.46
39.21
40.35
ASAP-CIR*
InternVL3.5-38B
29.42
30.03
32.67
33.69
Table 6: Robustness to the choice of MLLM parser on CIRCO. Both variants use the ViT-L/14 retrieval backbone and identical retrieval settings.
Figure 5: Ablation studies on CIRCO. (a) Effect of the atomic grounding temperature with holistic/atomic weights fixed to 0.95/0.05 . (b) Effect of holistic/atomic fusion weights with the grounding temperature fixed to τ=0.07 .
Figure 6: Qualitative examples on the CIRR validation split. Each row shows the reference image, the ground-truth target image, and the top-5 retrieved candidates. The first three rows are successful cases and the final two rows are failure cases. A green border marks the correctly retrieved target
Appendix figures & tables2 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 7: Operation-wise R@1 comparison between target-side static and dynamic composed query formulations under additional OpenCLIP backbones.
Figure 8: Target-anchored PCA projections under additional OpenCLIP backbones. For each operation type, 50 samples are randomly selected. Target-side static query features are generally closer to the target features than dynamic composed query features.