cs.CVOct 1, 2026

Beyond Domain-Level Adaptation: Margin-Oriented Semantic-Appearance Interaction Correction for Personalized Federated Vision-Language Models

Authors: Wentao Yue, Qingyu Mao, Tianyou Lai, Ahmed M. Abdelmoniem, Qilei Li

Abstract

Federated parameter-efficient fine-tuning enables distributed clients to adapt pretrained vision-language models without sharing raw data or updating the full backbone. Its effectiveness, however, is limited by domain heterogeneity across clients. Existing personalized methods separate globally shared knowledge from client-specific style, but they largely treat each domain as a class-agnostic transformation. We show that this abstraction is insufficient: the cross-domain displacement associated with a fixed domain varies across semantic classes, and only a subset of these class-domain residuals damages the image-text decision margin. We therefore propose Margin-Oriented Semantic-Appearance Interaction Correction (MOSAIC), which first constructs a decision-aware harmfulness score that measures whether a training-derived class-domain residual favors a competing text prototype over the true class. It then models fine-grained class-domain interactions with a low-rank residual adapter whose class factors and residual basis are globally shared while domain factors remain client-private. An image-conditioned gate further controls candidate-wise correction, and harmful-pair-aware reweighting prioritizes decision-relevant residuals during local optimization. Extensive experiments on Office31, OfficeHome, and DomainNet100 demonstrate that MOSAIC consistently improves macro-client top-1 accuracy across all evaluated domain-shift and joint domain-label-shift settings.

Figures & tables

Appendix figures & tables14 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Decoupled Training with Local Reinforcement Fine-Tuning in Federated Learning

    May 27, 2026Yuting Ma, Lechao Cheng, Xiaohua XuFederated LearningVision-Language Model Adaptation

  2. FedMPT: Federated Multi-label Prompt Tuning of Vision-Language Models

    May 27, 2026Xucong Wang, Pengkun Wang, Zhe Zhao +3Multi-Label ClassificationFedavg

  3. VisCoP: Visual Probing for Video Domain Adaptation of Vision Language Models

    Oct 15, 2025Dominick Reilly, Manish Kumar Govind, Le Xue +1Vision-Language Model AdaptationDomain Adaptation