cs.AI · 2610.01389 Copy arXiv ID · Oct 1, 2026 Save AiSearch: Interactive Multi-Modal Search with VLMs Authors: Ali Koksal , Mei Chee Leong , Vicky Sintunata , Ching Ling Chin , Wee Teck Fong
Organizations: Institute of Advanced Intelligence and Computing (IAIC), Agency for Science, Technology and Research (A*STAR), Singapore
Abstract Modern retrieval systems must both be automated and interactive, allowing users to search and refine results in real time. We present AiSearch, a flexible multimodal retrieval framework that leverages the zero shot capabilities of Vision Language Models (VLMs) for natural language search over images and videos. AiSearch supports interactive search refinement through user feedback to tailor results to the user's intent, and allows visual benchmarking across multiple VLMs, enabling users to select the most suitable model for their task.
Explore similar work Jun 30, 2026 · Ming Dai, Zhihong Lu, Jinjie Gu +5 Multimodal Search Agents Qwen3
Southeast University · Ant Group
Aug 31, 2026 · Jingyi He, Sanghwan Kim, Zeynep Akata Multiple Vision Tasks Recent Vision-Language Models
Technical University of Munich · Helmholtz Munich · Munich Center for Machine Learning
Aug 5, 2026 · Shengcao Cao, Tanmaya Shekhar Dabral, Zhongli Ding +6 Composed Image Retrieval
University of Illinois Urbana-Champaign · Google DeepMind · OpenAI. Work done at Google DeepMind.
Jun 30, 2026, cs.CV J/K move · Enter open · S save
Ming Dai, Zhihong Lu, Jinjie Gu, Jiedong Zhuang +4
Southeast University · Ant Group
We present SimpleSearch-VL, an efficient, reliable, and practical framework for multimodal agentic search. Its core idea is to improve the agent's own search-and-verification process rather than scaling data, tools, or auxiliary model components. For efficiency, Factorized Adaptive Rollout (FAR) improves sampling efficiency by forming more informative training groups while using redundant samples to mitigate long-tail latency and expose hard samples. For reliability, SimpleSearch-VL performs evidence-verified reasoning, explicitly using chain-of-thought verification to assess the relevance of retrieved visual and textual cues to the original context. For practicality, SimpleSearch-VL keeps a lightweight tool interface and performs webpage self-summary within the agent, requiring no additional external dependencies. With only 5K supervised tool-interleaved trajectories and 2K RL data, SimpleSearch-VL improves Qwen3-VL agentic baselines by 15.8 and 16.0 average points for the 8B and 30B-A3B variants, respectively. The SimpleSearch-VL-30B-A3B model further achieves performance competitive with agentic Gemini-3-Pro.