cs.CVOct 5, 2026

Semantic Capability Acquisition and Specialization During Vision-Language Model Fine-Tuning

Authors: Suguru Onda, Matthew Bailey, Ryan Farrell

Organizations: Brigham Young University · Carnegie Mellon University

Abstract

Fine-tuning vision-language models (VLMs) is typically evaluated at a single downstream checkpoint, obscuring whether a semantic capability was never acquired or emerged earlier and later declined during specialization. We ask how semantic capabilities are acquired, when they peak, how well they transfer, and what remains at deployment. We study these dynamics as a semantic capability trajectory, tracking identity- and attribute-based capabilities over training. We formulate a trajectory-based framework that separates capability acquisition, capability-specific optima, and later specialization, and introduce Structured Semantic Routing (SSR) to study how the representation of supervision shapes what is acquired. Across six pretrained backbones spanning DFN, MetaCLIP, and OpenAI CLIP, we show that fine-tuning can acquire semantic capability beyond the pretrained state, including gains observed on held-out evaluations. Unstructured name-and-attribute supervision produces strong name-and-attribute retrieval with comparatively weak name-free attribute-profile retrieval, whereas SSR yields substantially stronger name-free attribute-profile retrieval and is further strengthened by stochastic name-branch dropout. Different capabilities can peak at different stages, so a checkpoint selected by target class-name retrieval need not coincide with a transferable semantic optimum. Continued optimization can therefore preserve strong target class-name retrieval while reducing previously acquired transferable semantic capability. In a representative diagnostic study, this late specialization is consistent with reduced cross-modal semantic accessibility while substantial image-only class structure remains available.

Figures & tables

Appendix figures & tables63 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Text Capability Loss in Vision-Language Adaptation: An Attention-Sink Diagnosis

    Sep 1, 2026Minsik Choi, Geewook Kim, Young Geun KimVision-Language Model AdaptationDynamic Attention

  2. From Seeing to Thinking: Decoupling Perception and Reasoning Improves Post-Training of Vision-Language Models

    May 19, 2026Juncheng Wu, Hardy Chen, Haoqin Tu +6Recent Vision-Language ModelsVisual Reasoning

  3. Seeing and Solving Are Not Enough for Vision-Language Models

    Sep 27, 2026Ziheng Wang, Mingxuan Xie, Yilin Liu +3Multimodal QueryState-Tracking