Audio Representation Learning

Latest papers 141

All topics
CardsList
  1. Frequency-Aware Self-Supervised Music Representation Learning

    Jun 24, 2026Yicheng Gu, Junan Zhang, Jerry Li +2Audio Representation LearningFrequency-Domain Feature Learning

  2. PHAST-Net: Attention-Guided, Physics-Informed Network for Unified Estimation of Ideal Time-Frequency Representations

    Jun 22, 2026James M. Cozens, Simon J. GodsillAudio Representation LearningWavelet Methods

  3. STAR-VAE: Structured Topology-Aware Regularization for Audio Reconstruction and Generation

    Jun 22, 2026Huadai Liu, Wen Wang, Kaicheng Luo +3Variational AutoencodersText-to-Audio Generation

  4. Toward Open-Set Speaker Attribute Prediction with Keyword-Appended LLM Embeddings

    Jun 20, 2026Byoungjun So, Jaejun Lee, Kyogu LeeAudio Representation LearningSpeech Processing

  5. ProsoCodec: Prosody-Oriented Speech Codec for Voice Conversion

    Jun 20, 2026Jeongsoo Choi, Ji-Hoon Kim, Shujie Hu +1Audio Representation LearningNeural Audio Codecs

  6. CoughPhase-CLR: Designing an acoustics-informed foundation model for coughing sound classification

    Jun 19, 2026Marius Moldovan, Anton Batliner, Thomas M. Berghaus +2Contrastive LearningAudio Representation Learning

  7. LISE : Listenable Interpretable Speaker Embeddings

    Jun 19, 2026Xiaoliang Wu, Chongxin Gan, Ke Liu +2Audio Representation LearningSpeaker Verification

  8. Segment-Level Mandarin Chinese Speech-Based Cognitive Impairment Detection via an Autoencoder with Contrastive Learning

    Jun 18, 2026Yongqi Shao, Hong Huo, Flavio Bertini +2Audio Representation LearningRepresentation Learning

  9. Acoustic Prompting via Stage-wise Modulation for Few-Shot Learning in Audio Language Models

    Jun 14, 2026Hyebin Cho, Jaehyuk Jang, Changick Kim +1Few-Shot Audio ClassificationAudio Representation Learning

  10. Spectro-Temporal Interference Confounds Phase Encoding in Spatial Audio Foundation Models

    Jun 12, 2026Yuxuan Chen, Haoyuan Yu, Peize HeAudio Representation LearningDirection-of-Arrival Estimation

  11. Spatio-Temporal Audio Language Modeling for Dynamic Sound Sources

    Jun 12, 2026Oh Hyun-Bin, Kazuki Shimada, Yuhta Takida +6Audio-Language Model EvaluationAudio Representation Learning

  12. From Tokens to Faces: Investigating Discrete Speech Representations for 3D Facial Animation

    Jun 11, 2026Pedro Correa, Olivier Perrotin, Samir Sadok +2Audio Representation LearningNeural Audio Codecs

  13. Decoding Insect Song: A Multitask Semisupervised Orthoptera Bioacoustic Classifier

    Jun 11, 2026Olga Isupova, Danil Kuzin, Ella Browning +2Audio Representation LearningAudio Understanding

  14. From Physics to Representation: Audio Learning with Synthetic Pre-training via Procedural Generation

    Jun 11, 2026Fengrui Liu, Ruiyang Huang, Qijian Zheng +2Synthetic Data PretrainingAudio Representation Learning

  15. Dolph2Vec: Self-Supervised Representations of Dolphin Vocalizations

    Jun 10, 2026Chiara Semenzin, Faadil Mustun, Roberto Dessi +5Audio Representation LearningAudio Understanding

  16. Which Speech Representation Better Matches Text-Native Reasoning? A Study of Speech-Text Alignment on Frame Rate and Representation

    Jun 10, 2026Zhen Ye, Xu Tan, Yiming Li +10Spoken Language UnderstandingAudio Representation Learning

  17. SARA: A Dual-Stream VAE for High-Fidelity Speech Generation via Integrating Semantic and Acoustic Representations

    Jun 10, 2026Peijie Chen, Wenhao Guan, Weijie Wu +7Variational AutoencodersAudio Representation Learning

  18. Speaker Group Encoding in Self-supervised Speech Recognition Models

    Jun 9, 2026Felix Herron, Solange Rossato Alexandre Allauzen, Benoit Favre +1Audio Representation LearningDemographic Bias in ASR

  19. ViP-VL: Vietnamese Self-supervised Speech Pretraining Model with Vector-Quantization Learning

    Jun 9, 2026Khanh Le, Kiet Anh Hoang, Bao Nguyen +5Audio Representation LearningSpeech Processing

  20. Dual-Branch Gated Fusion for Open-Set Audio Deepfake Source Tracing

    Jun 8, 2026Awais Khan, Kutub Uddin, Khalid MalikAudio Representation LearningAudio Deepfake Detection

  21. Is Text All You Need? Text as a Universal Information Bottleneck for Speech LLMs

    Jun 8, 2026Ming-Hao Hsu, Yuxuan Hu, Shujie Liu +3Audio Representation LearningAudio Understanding

  22. Multi-View Speech Representation Learning for Parkinson's Disease Detection Using Context-guided Cross-modal Attention

    Jun 8, 2026George Theodosiou, Loukas Ilias, Dimitris AskounisAudio Representation LearningSpeech Processing

  23. Probing Token Spaces under Generator Shift in AI-Generated Music Detection

    Jun 7, 2026Joonyong Park, Jungwoo Kim, Junyoung Koh +1Audio Representation LearningAudio Deepfake Detection

  24. USAD 2.0: Scaling Representation Distillation for Universal Audio Understanding

    Jun 4, 2026Heng-Jui Chang, Alexander H. Liu, Saurabhchand Bhati +4Audio Representation LearningAudio Understanding

  25. F3-Tokenizer: Taming Audio Autoencoder Latents for Understanding and Generation

    Jun 4, 2026Dinghao Zhou, Xingchen Song, Di Wu +3Audio Representation LearningAutoencoders