cs.CVOct 7, 2026

FedSSMCoOp: SSM Encoders for light-weight Federated Prompt Learning for Few-shot Classification

Authors: Ankita Das, Ambarish Parthasarathy, Sumohana S. Channappayya, C. Krishna Mohan

Organizations: Department of Computer Science and Engineering Indian Institute of Technology Hyderabad · Department of Artificial Intelligence Indian Institute of Technology Hyderabad · Department of Electrical Engineering Indian Institute of Technology Hyderabad

Abstract

Vision-Language Models (VLMs) have shown strong performance across a wide range of downstream vision tasks, thanks to the complementary information contained in the respective domains. Despite the performance gains, most of these approaches rely on aligning these domains using the cosine similarity metric, which fails to capture token-level structure and cross-modal interactions prior to the classification stage. This is especially critical in biomedical applications under federated constraints, where data sharing is restricted, labeled data is scarce at each site, and it differs widely across institutions, leading to substantial statistical heterogeneity. To overcome this issue, we propose FedSSMCoOp, a federated few-shot image classification framework that enables multimodal learning while preserving data privacy. With the help of the SSM-based Vision Mamba and Cross Mamba blocks, and by optimizing only the soft-prompt and communication-prompt updates in the federated setting, the framework prioritizes both computation and performance. Importantly, this eliminates the need to use an external Large Language Model (LLM) for feature alignment. The framework is further trained and evaluated on various biomedical image datasets, and its performance is assessed. The proposed framework delivers stable performance relative to the baselines and is, on average, 1.96 times lighter. The corresponding script will be made available soon.

Figures & tables

Appendix figures & tables8 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. FM2^2: Unified Federated Foundation Models for Heterogeneous Multimodal Medical Imaging

    Jul 15, 2026Shengchao Chen, Ting ShuMultimodal Foundation ModelCross-Modality

  2. Multi-View Synergistic Learning with Vision-Language Adaption for Low-Resource Biomedical Image Classification

    Apr 27, 2026Xiaoliu Luo, Minxue Xiao, Ting Xie +5Vision-Language Model AdaptationMedical Image Classification

  3. FedMPT: Federated Multi-label Prompt Tuning of Vision-Language Models

    May 27, 2026Xucong Wang, Pengkun Wang, Zhe Zhao +3Multi-Label ClassificationFedavg