cs.CVJul 22, 2026

Diverse-Intent Multi-Turn Fashion Image Retrieval

Authors: Mingqiang TangHaokun WenMeng LiuYupeng HuWeili GuanXuemeng Song

Organizations: 1Southern University of Science and Technology, Shenzhen, China · 2Harbin Institute of Technology (Shenzhen), Shenzhen, China · 3Shandong University, Jinan, China · 4Shenzhen Loop Area Institute, Shenzhen, China

Abstract

Real-world fashion search involves interactive retrieval across multiple turns. However, existing multi-turn retrieval methods are built on a restrictive assumption that every interaction follows the same attribute-editing paradigm, leaving heterogeneous intent transitions unexplored. Moreover, existing approaches often rely on textification to bridge multimodal queries and visual retrieval, which may lose fine-grained visual cues. To address these gaps, we introduce DIM-Fashion, a benchmark of 26K multi-turn sessions constructed from 13 fashion retrieval datasets across 7 tasks, featuring diverse intent transitions and rollback behaviors. We further propose FashionAM, an MLLM-VLP framework that directly aligns multimodal conversational queries with a fashion-oriented gallery embedding space, avoiding intermediate textification. Extensive experiments demonstrate the effectiveness of FashionAM over existing approaches. The dataset and code will be made publicly available upon acceptance.

Explore similar work

CardsList