Audio Understanding

Latest papers 169

All topics
CardsList
  1. Learning When to Think While Listening in Large Audio-Language Models

    May 26, 2026Zhiyuan Song, Weici Zhao, Yang Xiao +3Audio QARL for Language Model Reasoning

  2. Beyond Binary: Speech Representations Across the Cognitive Score Hierarchy

    May 26, 2026Serli Kopar, Roshan Prakash Rane, Christian Mychajliw +6Speech ProcessingAudio Understanding

  3. Why Can't They Remember? Uncovering Representation and Retrieval Bottlenecks in Multi-Turn Acoustic Memory

    May 26, 2026Yang Xiao, Siyi Wang, Han Yin +4Audio-Language Model EvaluationMemorization in Language Models

  4. O-MARC: Omni Memory-Augmented Compression Distillation for Efficient Video Understanding

    May 26, 2026Peiran Wu, Yunze Liu, Chi-Hao Wu +2Audio-Visual UnderstandingEfficient Multimodal Inference

  5. Thaka at KSAA-2026 Task 2: Regularized Fine-Tuning for Arabic Speech Diacritization

    May 25, 2026Meshal Alamr, Hassan Alqaeri, Abdullah AldahlawiFine-TuningArabic NLP

  6. PitchBench: Measuring Pitch Hearing in Audio-Language Models

    May 25, 2026Milan Liessens Dujardin, Song-Ze Yu, Craver Corbyn Thomas-Smith +2Audio-Language Model EvaluationAudio Understanding

  7. Zero-Shot Parkinson's Disease Detection from Speech: Comparing Large Audio and Language Models

    May 24, 2026Muhammad Ashad Kabir, Sirajam MuniraAudio-Language Model EvaluationZero-Shot Learning

  8. Evaluating the Temporal Detection Capability of Integrated Gradients Applied on Sound Classifier

    May 22, 2026Martynas Dumpis, Tuomas VirtanenGradient-Based AttributionIntegrated Gradients

  9. SpeakerLLM: A Speaker-Specialized Audio-LLM for Speaker Understanding and Verification Reasoning

    May 14, 2026KiHyun Nam, Jungwoo Heo, Siu Bae +2Audio Representation LearningSpeaker Verification

  10. AudioMosaic: Contrastive Masked Audio Representation Learning

    May 14, 2026Hanxun Huang, Qizhou Wang, Xingjun Ma +3Audio Representation LearningSelf-Supervised Pre-Training

  11. FSD50K-Solo: Automated Curation of Single-Source Sound Events

    May 13, 2026Ningyuan Yang, Sile Yin, Li-Chia Yang +4Training Data CurationAudio Understanding

  12. SpurAudio: A Benchmark for Studying Shortcut Learning in Few-Shot Audio Classification

    May 13, 2026Giries Abu Ayoub, Morad Tukan, Loay MualemFew-Shot Audio ClassificationAudio Understanding

  13. NAACA: Training-Free NeuroAuditory Attentive Cognitive Architecture with Oscillatory Working Memory for Salience-Driven Attention Gating

    May 13, 2026Zhongju Yuan, Geraint Wiggins, Dick BotteldoorenAudio UnderstandingAudio-Language Models

  14. When Vision Speaks for Sound

    May 13, 2026Xiaofei Wen, Wenjie Jacky Mo, Xingyu Fu +6Audio-Language Model EvaluationAudio-Visual Understanding

  15. A Semi-Supervised Framework for Speech Confidence Detection using Whisper

    May 12, 2026Adam Wynn, Jingyun WangAudio UnderstandingPseudo-Labeling

  16. APEX: Audio Prototype EXplanations for Classification Tasks

    May 11, 2026Piotr Kawa, Kornel Howil, Piotr Borycki +3Explainable Artificial IntelligenceAudio Understanding

  17. Omni-DeepSearch: A Benchmark for Audio-Driven Omni-Modal Deep Search

    May 9, 2026Tao Yu, yiming ding, Shenghua Chai +16Multi-Hop QAMultimodal IR

  18. WavCube: Unifying Speech Representation for Understanding and Generation via Semantic-Acoustic Joint Modeling

    May 7, 2026Guanrou Yang, Tian Tan, Qian Chen +12Audio Representation LearningTTS Synthesis

  19. MultiLinguahah : A New Unsupervised Multilingual Acoustic Laughter Segmentation Method

    May 7, 2026Sofia Callejas, Nahuel Gomez, Catherine Pelachaud +2Audio Understanding

  20. Minimizing Modality Gap from the Input Side: Your Speech LLM Can Be a Prosody-Aware Text LLM

    May 7, 2026Wenqian Cui, Xiao-Hui Li, Daxin Tan +2Audio UnderstandingSpeech Prosody

  21. Bangla-WhisperDiar: Fine-Tuning Whisper and PyAnnote for Bangla Long-Form Speech Recognition and Speaker Diarization

    May 6, 2026Mohammed Aman Bhuiyan, Md Sazzad Hossain Adib, Samiul Basir Bhuiyan +4Audio UnderstandingAutomatic Speech Recognition

  22. VocalParse: Towards Unified and Scalable Singing Voice Transcription with Large Audio Language Models

    May 6, 2026Yukun Chen, Tianrui Wang, Zhaoxi Mu +2Audio UnderstandingMusic Information Retrieval

  23. JASTIN: Aligning LLMs for Zero-Shot Audio and Speech Evaluation via Natural Language Instructions

    May 6, 2026Leying Zhang, Bowen Shi, Haibin Wu +2Audio-Language Model EvaluationZero-Shot Learning

  24. PHALAR: Phasors for Learned Musical Audio Representations

    May 5, 2026Davide Marincione, Michele Mancusi, Giorgio Strano +4Audio Representation LearningRepresentation Learning

  25. ReasonAudio: A Benchmark for Evaluating Reasoning Beyond Matching in Text-Audio Retrieval

    May 5, 2026Honglei Zhang, Yuting Chen, Chenpeng Hu +2Audio ReasoningAudio Understanding