Sep 21, 2026 · cs.AIJ/K move · Enter open · S save
Minda Zhao, Fangyu Hu, Yan Luo, Yutong Yang+7
Harvard AI and Robotics Lab, Schepens Eye Research Institute of Massachusetts Eye and Ear, Harvard Medical School, Boston, MA, USA · Department of Ophthalmology, Massachusetts Eye and Ear, Harvard Medical School, Boston, MA, USA · Department of Ophthalmology, National Taiwan University Hospital, Taipei City, Taiwan · MIT Media Lab, Massachusetts Institute of Technology, Cambridge, MA, USA · Department of Computer Science, McCormick School of Engineering, Northwestern University, Evanston, IL, USA
Medical image interpretation is central to diagnosis and care, yet adapting general-purpose multimodal large language models (MLLMs) often requires resource-intensive domain-specific fine-tuning. Here we introduce representation-guided in-context learning (RG-ICL), a training-free inference framework that retrieves query-aligned demonstrations using frozen encoders, without task-specific parameter updates. Across eight datasets spanning histopathology, radiology and retinal fundoscopy, RG-ICL improved classification (mean gain 20 percentage points) and visual question answering (VQA) (mean gain 13 percentage points) over no-context and conventional ICL, approaching or exceeding training-based comparators. Which cases were retrieved mattered more than how many: 6 query-aligned cases outperformed up to 32 randomly selected ones, whereas fixed or random cases often reduced accuracy below baseline. For VQA, aligning reference cases with both image content and question intent produced further gains. These findings indicate that for medical image interpretation, curating which reference cases an MLLM sees is a practical alternative to retraining it.