cs.CVJun 18, 2026

Exploring Multi-Modal Large Language Models and Two-Stage Fine-Tuning for Fashion Image Retrieval

Authors: Nguyen Cao HoangHoang Bui LeNam Vo HoangTrung-Nghia Le

Organizations: University of Science, VNU-HCM, Ho Chi Minh, Vietnam · Vietnam National University, Ho Chi Minh, Vietnam

Abstract

Composed image retrieval retrieves a target image using a composed query of a reference image and a modified text description. In the fashion domain, this task requires understanding subtle attribute variations such as color, pattern, and texture. However, existing approaches face limitations due to scarce annotated data and simplistic negative sampling. We propose a novel framework that integrates a multi-modal large language model (LLaVA) to generate attribute-aware triplets and introduces a two-stage fine-tuning strategy to enhance contrastive learning. We leverage pretrained vision-language models, such as CLIP-ViT/B32, to generate and concatenate sentence-level prompts with the relative caption and to scale the number of negatives using static representations. Experimental results demonstrate enhanced compositional reasoning and improved fine-grained retrieval behavior, underscoring the feasibility and potential of the proposed framework for fashion retrieval.

Explore similar work

CardsList
  1. Diverse-Intent Multi-Turn Fashion Image Retrieval

    Jul 22, 2026Mingqiang Tang, Haokun Wen, Meng Liu +3GarmentsMultimodal Query