cs.SDOct 8, 2026

A computational framework for the acoustic characterization of Suikinkutsu (water harp cave)

Authors: Paul Haimes, Alarith Uhde, Daniel Moritz Marutschke

Organizations: Ritsumeikan University

Abstract

Suikinkutsu are partially concealed acoustic devices traditionally installed adjacent to stone water basins (tsukubai) in Japanese temple gardens. While previous research has examined their physical acoustics and cultural significance, computational methods for documenting and comparing their acoustic characteristics remain underexplored. This exploratory study introduces a computational framework for analyzing suikinkutsu using spectral and temporal audio features extracted from field recordings. Rather than estimating the physical geometry of individual suikinkutsu, the framework characterizes their acoustic profiles using a multidimensional set of computational descriptors, enabling systematic comparisons between recordings across different sites and installations. Audio recordings from nine suikinkutsu located at temples, shrines, and gardens in the Kansai region of Japan were analyzed using a feature-extraction pipeline that incorporates Fast Fourier Transform (FFT), Mel-Frequency Cepstral Coefficients (MFCCs), Root Mean Square (RMS) energy, spectral centroid, spectral bandwidth, and spectral contrast. The extracted features indicate that, although the suikinkutsu share broadly similar spectral and timbral characteristics, each exhibits a measurably distinct acoustic profile; a supplementary analysis of within-site and between-site variability suggests these profiles reflect a bounded range of characteristic resonant behaviour rather than a fixed acoustic signature. The proposed framework provides one of the first computational approaches for systematically characterizing and comparing suikinkutsu, contributing a practical foundation for future research and the digital preservation of this distinctive form of Japanese environmental design.

Figures & tables

Explore similar work

Jun 12, 2026cs.LG

Beyond task performance: Decoding bioacoustic embeddings with speech features

Pretrained audio embeddings are standard in bioacoustics, yet little is known about which acoustic features these models encode, nor which are useful for a given task. This hinders transparency and limits extension to rare species or data-scarce domains. Here we reveal which speech-like features are encoded in bioacoustic representations. Using the 88~eGeMAPS features across six taxonomic groups, we apply linear and nonlinear regression probes to quantify which acoustic properties each model captures. Results confirm a ``no free lunch'' pattern: no single model captures the full feature space. A concatenated embedding achieves the highest performance, suggesting complementary acoustic space coverage across models. Loudness features are best encoded (R2=0.76R^2 = 0.76) while F0 is hardest to recover (R2=0.33R^2 = 0.33). By cross-referencing recoverability with per-species feature salience (NMI), we derive data-driven model selection guidance for bioacoustics.
Jul 3, 2026eess.AS

An Intervention-Based Framework for Shortcut Diagnosis in Spoofing Countermeasures

While deepfake audio detection systems achieve high performance in controlled benchmarks, their reliability often diminishes in the wild. Prior work shows that dataset-specific artifacts contribute to this gap. Yet, systematic tools to identify which acoustic properties a model exploits as shortcuts remain limited. We propose an intervention-based diagnostic framework, grounded in a directed graphical model, that formally distinguishes confound-driven shortcut dependencies from legitimate domain shift. We operationalise this through controlled acoustic perturbations targeting non-speech structure, spectral content, and signal energy, complemented by corpus-level distributional analysis. Evaluating XLS-R-300M with RawGAT-ST across ASVspoof challenges datasets, we quantify model sensitivity to specific intervention types. Results reveal that non-speech interventions produce the largest performance shifts, confirming non-speech intervals as a dominant shortcut.
May 5, 2026cs.CR

DECKER: Domain-invariant Embedding for Cross-Keyboard Extraction and Recognition

Acoustic side-channel attacks (ASCA) on keyboards pose a significant security risk, as keystrokes can be inferred from typing acoustics, revealing sensitive information. Prior ASCA studies are limited by small-scale datasets with restricted diversity in users, keyboards, and environments, constraining analysis across devices, microphones, and noise conditions. We introduce HEAR, a dataset designed to study ASCA along three axes: keyboard generalization, noise adaptation, and user bias. HEAR contains recordings from 53 participants using 37 laptop keyboards, collected in three realistic settings: (1) external microphone capture, (2) device microphone capture without network noise, and (3) VoIP-based streaming capture. This enables controlled evaluation across users, keyboards, and environments. On HEAR, we establish an ASCA benchmark spanning conventional features and pre-trained representations from raw audio and spectrograms in unimodal and multimodal settings. We propose DECKER, a domain-invariant keystroke inference framework with four stages: (1) Keyboard Signature Normalization to reduce device coloration, (2) domain-adversarial disentanglement to suppress keyboard identity, (3) supervised cross-keyboard contrastive alignment to enforce key consistency, and (4) Acoustic Style Randomization to synthesize unseen keyboard responses. We further explore sentence-level inference using an LLM-based post-processing layer to refine keystroke sequences via linguistic context. Results on HEAR show DECKER improves keystroke identification over strong baselines, particularly in cross-keyboard and cross-user settings, with further gains from language-model rectification. These findings highlight that ASCA remains effective across diverse users, devices, and noisy environments, underscoring its practical security risk.