eess.ASDec 2, 2024

TACO: Training-free Sound Prompted Segmentation via Semantically Constrained Audio-visual CO-factorization

Authors: Hugo Malard, Michel Olvera, Stephane Lathuiliere, Slim Essid

Organizations: LTCI, Télécom Paris, Institut Polytechnique de Paris

Abstract

Large-scale pre-trained audio and image models demonstrate an unprecedented degree of generalization, making them suitable for a wide range of applications. Here, we tackle the specific task of sound-prompted segmentation, aiming to segment image regions corresponding to objects heard in an audio signal. Most existing approaches tackle this problem by fine-tuning pre-trained models or by training additional modules specifically for the task. We adopt a different strategy: we introduce a training-free approach that leverages Non-negative Matrix Factorization (NMF) to co-factorize audio and visual features from pre-trained models so as to reveal shared interpretable concepts. These concepts are passed on to an open-vocabulary segmentation model for precise segmentation maps. By using frozen pre-trained models, our method achieves high generalization and establishes state-of-the-art performance in unsupervised sound-prompted segmentation, significantly surpassing previous unsupervised methods.

Figures & tables

Appendix figures & tables11 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Bagpiper: Solving Open-Ended Audio Tasks via Rich Captions

    Feb 5, 2026Jinchuan Tian, Haoran Wang, Bo-Hao Su +14Audio UnderstandingNeural Audio

  2. Training-Free Generalized Few-Shot Segmentation through Open-Vocabulary Semantic Arbitration

    Jun 8, 2026Silas Kwabla Gah, Ebenezer OwusuFew-Shot SegmentationOpen-Vocabulary

  3. UniAudio-Token: Empowering Semantic Speech Tokenizers with General Audio Perception

    May 29, 2026Yuhan Song, Linhao Zhang, Aiwei Liu +6Audio TokenizersAcoustic Latent Space