eess.ASDec 19, 2025

Review of MEMS Transducers for Audio Applications

Authors: Nils WittekAnton MelnikovBert KaiserAndré Zimmermann

Organizations: University of Stuttgart, Institute for Micro Integration (IFM), 70569 Stuttgart, Germany · Bosch Sensortec GmbH, Robert-Bosch-Ring 1, 01109 Dresden, Germany · Hahn-Schickard, 70569 Stuttgart, Germany

Abstract

Microelectromechanical systems (MEMS) speakers are compact, scalable alternatives to traditional voice coil speakers, promising improved sound quality through precise semiconductor manufacturing. This review provides an overview of the research landscape, covering baseband-displacement, ultrasound-based, and thermoacoustic sound generation concepts, classifying MEMS speakers by their actuation principles as electrodynamic, piezoelectric, or electrostatic devices. A comparative analysis of performance indicators from 1990 to 2026 highlights the dominance of piezoelectric MEMS with baseband displacement, focusing on miniaturization and efficiency. The review outlines upcoming research challenges and identifies potential candidates for achieving full-spectrum audio performance. A focus on innovative approaches could lead to widespread adoption of MEMS-only speakers.

Explore similar work

Jun 19, 2026cs.RO

Membrane-based Acoustic Microrobots

Acoustic microrobots have emerged as a promising frontier for targeted drug delivery and minimally invasive medicine due to their high-power density and biocompatibility. Despite wide-ranging designs, conventional acoustic microrobots mostly rely on air microbubbles trapped within confined microcavities within the robot body, which suffer from limited operational longevity due to rapid gas dissolution and resultant shifts in resonance frequency. In this paper, we propose a robust, membrane-based acoustic microrobot that overcomes these limitations by employing a thin flexible Polydimethylsiloxane (PDMS) membrane bonded over confined microcavities for microstreaming. The introduced design physically prevents gas diffusion, ensuring stable performance over extended periods at high actuation voltages. We systematically characterized the membrane-based acoustic actuator longevity, demonstrating consistent streaming and propulsion for over 24 hours of continuous operation. In addition, by embedding magnetic microparticles into the structural body, these actuators were successfully employed as microswimmers with directional control using low-intensity (2 mT) external magnetic fields. Finally, we demonstrate the scalability of the proposed design architecture down to ~100 um. This membrane-based approach establishes a reliable framework for the development of high-endurance acoustic microactuators and microrobots capable of performing long-term tasks.
Fatih Kocabas, Cemal Polat Avdar, Prithvi Venkatesh +1
Sep 9, 2026eess.AS

Pushing the Boundaries of Streaming Multi-Speaker ASR: A Systematic Study of Architectural Trade-offs

Streaming multi-speaker ASR is a challenging task that must balance accuracy, latency, and efficiency while handling overlapping speech and maintaining coherent long-context modeling over extended conversations in an online fashion. We present a unified framework that categorizes streaming multi-speaker ASR into four architectural strategies based on how diarization and ASR are integrated. Using a shared pair of open-source streaming ASR and diarization models as a common foundation, we derive four multi-speaker ASR systems that differ in whether they employ multiple model instances, fine-tuning, or both. We evaluate these systems across multi-speaker accuracy, single-speaker accuracy degradation, memory footprint, and training complexity. Through this systematic architectural analysis, we clarify the design space for streaming multi-speaker ASR and provide practical guidance for selecting the most suitable approach under diverse deployment constraints.
Taejin Park, Ivan Medennikov, Kunal Dhawan +3
Jul 21, 2026cs.SD

CAPS: A Cascaded Reconstruction Model to Power Saving in Hearables Using Sub-Nyquist Sampling with Bandwidth Extension

Hearables are wearable computers worn on the ear. Bone conduction microphones are used with air conduction microphones in hearables for multimodal speech enhancement in noisy conditions. Despite this potential, current models largely fail to explore how jointly reducing sampling bit resolution and sampling frequency in analog-to-digital converters (ADCs) of hearables impacts both power usage and audio quality. Furthermore, current frameworks cannot do sub-Nyquist sampling in hearables because they lack a method to reconstruct wideband signals from narrowband components. We therefore propose CAPS, which (i) intentionally employs sub-Nyquist sampling and low bit resolution in ADCs, achieving a 3.3x reduction in power consumption in hearables, and (ii) supports streaming operation on mobile platforms with an inference time of 1.36 ms and a memory footprint of 11.04 MB. CAPS ensures robust speech intelligibility in real-world settings, bridging the gap between efficiency and power savings.
Tarikul Islam Tamiti, Sajid Fardin Dipto, Luke Baja-Ricketts +2