Composed image retrieval (CIR) aims to retrieve a desired target image from a query consisting of a reference image and a modification text. This task exhibits an unusual representational asymmetry: the modification text specifies a transition from the reference state, whereas retrieval candidates depict completed target states. This creates a representation mismatch for zero-shot methods that query pretrained vision-language spaces directly with transformation-oriented language. We study this mismatch and reformulate zero-shot composed image retrieval as target-state reconstruction followed by retrieval. We instantiate this formulation with ASAP-CIR, a training-free framework that reconstructs a static target representation using a frozen multimodal large language model (MLLM). The representation combines multiple holistic descriptions with a variable set of importance-weighted atomic semantics, thereby preserving both overall target identity and fine-grained visual constraints. Retrieval then integrates holistic state alignment, atomic constraint grounding, and calibrated target-state evidence aggregation. A controlled text-only diagnostic shows that target-side static query formulations achieve more reliable retrieval than dynamic composed query formulations, particularly when source-state semantics must be suppressed or transformed. Experiments on FashionIQ, CIRR, and CIRCO further characterize the effectiveness and limitations of this representation principle, with the clearest gains on the multi-target CIRCO benchmark. These results show that how composed intent is represented before retrieval is a consequential design choice, distinct from the choice of retrieval backbone itself.
Figures & tables
Figure 1: Motivation of ASAP-CIR. The left panels compare three query formulation paradigms: (a) pseudo-word inversion directly combines an inverted visual token with dynamic modification text; (b) MLLM-based reasoning generates a target description without explicitly constraining it to static visual semantics; and (c) ASAP-CIR reconstructs the composed query as complementary holistic descriptions and importance-weighted atomic semantics for multi-level retrieval. Panel (d) visualizes the phenomenon observed in our diagnostic experiments: target-side static query embeddings are more closely aligned with the target embedding region, whereas dynamic composed query embeddings are more dispersed and frequently deviate from the intended target state.
Figure 2: Overview of the proposed ASAP-CIR framework. Given a composed query, the MLLM parser performs target-state semantic reconstruction and produces complementary holistic descriptions and importance-weighted atomic semantics. In the figure, these outputs are labeled “Holistic Descriptions” and “Weighted Atomic Semantics,” respectively. The holistic descriptions are encoded independently and combined by average pooling before being aligned with global image features to obtain the holistic evidence score. The atomic semantics are encoded independently and grounded against candidate patch features through patch-wise softmax assignment and importance-weighted aggregation, producing the atomic evidence score. Finally, query-wise z-score normalization calibrates both evidence scores, and weighted linear fusion produces the final retrieval score and ranking.
Backbone
Method
Venue
T-Free
Dress
Shirt
Toptee
Average
R@10
R@50
R@10
R@50
R@10
R@50
R@10
R@50
ViT-B/32
SEARLE
ICCV’23
N
18.54
39.51
24.44
41.61
25.70
46.46
22.89
42.53
FlowCIR
ECCV’26
N
24.80
41.80
19.10
40.10
26.40
47.80
23.40
43.20
CIReVL
ICLR’24
Y
25.29
46.36
28.36
47.84
31.21
53.85
28.29
49.35
LDRE
SIGIR’24
Y
19.97
41.84
27.38
46.27
27.07
48.78
24.81
45.63
AutoCIR
KDD’25
Y
26.52
46.36
32.43
51.67
33.96
56.09
30.97
51.37
Table 1: FashionIQ validation results. The best and second-best results within each backbone group are shown in bold and underlined, respectively.
Backbone
Method
Venue
T-Free
R@1
R@5
R@10
R@50
ViT-B/32
SEARLE
ICCV’23
N
24.00
53.42
66.82
89.78
TT-RLDR
AAAI’26
N
27.91
–
71.98
91.89
FlowCIR
ECCV’26
N
25.50
56.50
69.80
–
CIReVL
ICLR’24
Y
23.94
52.51
66.00
86.95
LDRE
SIGIR’24
Y
25.69
55.13
69.04
89.90
OSrCIR
CVPR’25
Y
25.42
54.54
68.19
–
Table 2: CIRR test1 results. The best and second-best results within each backbone group are shown in bold and underlined, respectively.
Backbone
Method
Venue
T-Free
mAP@5
mAP@10
mAP@25
mAP@50
ViT-L/14
Pic2Word
CVPR’23
N
8.72
9.51
10.64
11.29
SEARLE
ICCV’23
N
11.68
12.73
14.33
15.12
LinCIR
CVPR’24
N
12.62
13.40
14.81
15.69
DiffComp
CVPR’26
N
16.19
17.32
19.24
20.30
FlowCIR
ECCV’26
N
14.90
15.70
17.30
18.20
CIReVL
ICLR’24
Y
18.57
19.01
20.89
21.80
Table 3: CIRCO test results. The best and second-best results within each backbone group are shown in bold and underlined, respectively.
Backbone
Static R@1
Dynamic R@1
Gap
ViT-B/32
81.52
29.69
51.83
ViT-L/14
83.23
29.00
54.22
ViT-G/14
86.21
32.59
53.61
Table 4: Controlled text-only retrieval results for target-side static and dynamic composed query formulations under different OpenCLIP backbones. Gap denotes the R@1 difference between the two formulations.
Figure 3: R@1 comparison between target-side static and dynamic composed query formulations under the default OpenCLIP ViT-L/14 text encoder. The synthetic evaluation set contains 4,096 Qwen3-32B generated samples across eight dynamic semantic categories.
Figure 4: Target-anchored principal component analysis (PCA) projection of CLIP text features under the default OpenCLIP ViT-L/14 text encoder. For each of the eight operation types, 50 query samples are randomly selected. Squares denote target-side static query features, crosses denote dynamic composed query features, and triangles denote target features.
Matching evidence
Importance weighting
mAP@5
mAP@10
mAP@25
mAP@50
Holistic only
–
35.19
36.14
38.88
40.03
Atomic only
×
11.96
12.61
13.92
14.66
Atomic only
✓
13.19
13.84
15.26
16.07
Holistic + atomic
×
35.40
36.39
39.10
40.25
Holistic + atomic
✓
35.49
36.46
39.21
40.35
Table 5: Component ablation of holistic state alignment, atomic constraint grounding, and importance weighting on CIRCO with ViT-L/14.
Variant
MLLM parser
mAP@5
mAP@10
mAP@25
mAP@50
ASAP-CIR
Qwen3.6-27B
35.49
36.46
39.21
40.35
ASAP-CIR*
InternVL3.5-38B
29.42
30.03
32.67
33.69
Table 6: Robustness to the choice of MLLM parser on CIRCO. Both variants use the ViT-L/14 retrieval backbone and identical retrieval settings.
Figure 5: Ablation studies on CIRCO. (a) Effect of the atomic grounding temperature with holistic/atomic weights fixed to 0.95/0.05 . (b) Effect of holistic/atomic fusion weights with the grounding temperature fixed to τ=0.07 .
Figure 6: Qualitative examples on the CIRR validation split. Each row shows the reference image, the ground-truth target image, and the top-5 retrieved candidates. The first three rows are successful cases and the final two rows are failure cases. A green border marks the correctly retrieved target
Appendix figures & tables2 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 7: Operation-wise R@1 comparison between target-side static and dynamic composed query formulations under additional OpenCLIP backbones.
Figure 8: Target-anchored PCA projections under additional OpenCLIP backbones. For each operation type, 50 samples are randomly selected. Target-side static query features are generally closer to the target features than dynamic composed query features.
Composed Image Retrieval (CIR) aims to retrieve target images by integrating a reference image with a corresponding modification text. CIR requires jointly considering the explicit semantics specified in the query and the implicit semantics embedded within its bi-modal composition. Recent training-free Zero-Shot CIR (ZS-CIR) methods leverage Multimodal Large Language Models (MLLMs) to generate detailed target descriptions, converting the implicit information into explicit textual expressions. However, these methods rely heavily on the textual modality and fail to capture the fuzzy retrieval nature that requires considering diverse combinations of candidates. This leads to reduced diversity and accuracy in retrieval results. To address this limitation, we propose a novel training-free method, Geodesic Mixup-based Implicit semantic eXpansion and Explicit semantic Re-ranking for ZS-CIR (G-MIXER). G-MIXER constructs composed query features that reflect the implicit semantics of reference image-text pairs through geodesic mixup over a range of mixup ratios, and builds a diverse candidate set. The generated candidates are then re-ranked using explicit semantics derived from MLLMs, improving both retrieval diversity and accuracy. Our proposed G-MIXER achieves state-of-the-art performance across multiple ZS-CIR benchmarks, effectively handling both implicit and explicit semantics without additional training. Our code will be available at https://github.com/maya0395/gmixer.
Composed image retrieval requires identifying a target image from a gallery by integrating a reference image with a textual modification instruction. In a training-free zero-shot setting, this task relies on constructing a retrieval-oriented textual query within a frozen vision--language embedding space at inference time. Existing approaches predominantly rely on a single-pass generation strategy that fuses the reference context and modification text into a unified description. This strategy makes it difficult to detect or correct semantic distortions and omissions during generation. Consequently, the preservation of reference attributes and the integration of textual requirements interfere with each other, which degrades retrieval precision. To address these challenges, we introduce PEC-CIR, a training-free framework that structures query construction as a multi-stage reasoning pipeline. The framework operates through a Planner--Executor--Critic architecture where the Planner extracts explicit constraints, the Executor generates multiple candidate target descriptions, and the Critic evaluates these candidates based on constraint compliance. By reframing query construction as a staged inference process instead of a single-pass output, PEC-CIR reduces the propagation of generative errors by explicitly evaluating candidate queries before retrieval, thereby improving retrieval stability.
Gunho Jung, Jeong-Woo Park, Seon Bin Kim +1
Department of Artificial Intelligence, Korea University, Seoul, 02841, Republic of Korea
Zero-shot composed image retrieval (ZS-CIR) aims to retrieve a target image from a multimodal query consisting of a reference image and an edit text describing the desired modification. Recent ZS-CIR studies have relied on projection-based methods that map a reference image into pseudo-word tokens in the text embedding space. However, such methods require additional projection and re-encoding steps, increasing training complexity, reducing efficiency, and introducing a discrepancy between training and inference. In this paper, we propose DiCE-CIR, a direct composition learning method that predicts composed query representations by directly composing a reference image and an edit text. To enable scalable training without manually annotated triplets, we automatically construct compositional training samples from large-scale image-caption pairs using a large language model. Based on these samples, we train a lightweight composition module with objectives that promote alignment with the target, edit-consistent semantic transformation, and retrieval discriminability. We conduct extensive experiments on ZS-CIR benchmarks and show that DiCE-CIR achieves state-of-the-art performance on CIRCO and competitive performance on CIRR while maintaining high computational efficiency.
Gwang-Ho Na, Ho-Joong Kim, Seong-Whan Lee
Department of Artificial Intelligence, Korea University, Anam-dong, Seongbuk-ku, Seoul 02841, Korea