cs.LGNov 21, 2025

PrismSSL: One Interface, Many Modalities; A Single-Interface Library for Multimodal Self-Supervised Learning

Authors: Melika Shirian, Kianoosh Vadaei, Kian Majlessi, Audrina Ebrahimi, Peyman Adibi, Hossein Karshenas

Organizations: dept. Computer Engineering University of Isfahan Isfahan, Iran · dept. Computer Engineering University of Texas at Dallas Isfahan, Iran · dept. Artificial Intelligence University of Isfahan Isfahan, Iran

Abstract

We present PrismSSL, a Python library that unifies state-of-the-art self-supervised learning (SSL) methods across audio, vision, graphs, and cross-modal settings in a single, modular codebase. The goal of the demo is to show how researchers and practitioners can: (i) install, configure, and run pretext training with a few lines of code; (ii) reproduce compact benchmarks; and (iii) extend the framework with new modalities or methods through clean trainer and dataset abstractions. PrismSSL is packaged on PyPI, released under the MIT license, integrates tightly with HuggingFace Transformers, and provides quality-of-life features such as distributed training in PyTorch, Optuna-based hyperparameter search, LoRA fine-tuning for Transformer backbones, animated embedding visualizations for sanity checks, Weights & Biases logging, and colorful, structured terminal logs for improved usability and clarity. In addition, PrismSSL offers a graphical dashboard - built with Flask and standard web technologies - that enables users to configure and launch training pipelines with minimal coding. The artifact (code and data recipes) will be publicly available and reproducible.

Explore similar work

CardsList
  1. FairSSL: Fair Multimodal Self-Supervised Learning

    Aug 22, 2025Jiaee Cheong, Abtin Mogharabin, Paul Liang +2Multimodal LearningSelf-Supervised Learning

  2. SSL-R1: Self-Supervised Visual Reinforcement Post-Training for Multimodal Large Language Models

    Apr 22, 2026Jiahao Xie, Alessio Tonioni, Nathalie Rauschmayr +2Multimodal Large Language ModelsRecent Vision-Language Models

  3. InsideSSL: Understanding Self-Supervised Speech Representations using a Model-Centric Perspective

    Jul 7, 2026Samir Sadok, Xavier Alameda-PinedaSelf-Supervised Speech ModelsWav2Vec