SincDPNet: Interpretable Raw-Waveform Bathroom Activity Recognition for Assistive Living
Authors: Debolina Chowdhury, Suman Samui, Sujoy Saha
Organizations: Department of Computer Science and Engineering, National Institute of Technology Durgapur, Durgapur 713209, West Bengal, India · Department of Electronics and Communication Engineering, National Institute of Technology Durgapur, Durgapur 713209, West Bengal, India
Bathroom acoustic-event recognition can support ambient assisted living in settings where continuous video monitoring is undesirable. However, practical deployment requires models that are compact, interpretable, and robust to changes in the recording environment. This work introduces \dataset{}, a seven-class bathroom acoustic-event dataset containing 21{,}387 annotated clips recorded across five environments, and proposes SincDPNet, a compact raw-waveform classifier with a learnable sinc filter bank followed by a depthwise-separable convolutional body. Each sinc filter is controlled by two frequency parameters, allowing the learned passbands to be inspected directly in hertz while keeping the front end small. To reduce room-specific leakage, recording sessions and environments are separated before overlapping windows are assigned to the training, validation, and test partitions. We further use multi-objective Bayesian optimization as a design tool to examine the validation performance--model-size trade-off across 24 configurations. The selected designs span different operating points: the best-performing model achieves 80.2% accuracy and 0.760 macro-F1 with 14{,}040 parameters, while the compact Nf=25 configuration uses only 2{,}848 parameters and achieves 75.7% accuracy, 0.661 macro-F1, and 0.716 MCC on the held-out environment. Analysis of the learned filters and confusion patterns shows that spectral overlap contributes to confusion among water-related events, while the \textit{Door}/\textit{Walker/Crutch} errors also reflect similarities in their transient temporal structure.
Figures & tables
Dataset
Domain
Bathroom
Public
Non-target class
Raw-waveform baselines
ESC-50 [ 29 ]
Environmental
✗
✓
✗
✗
UrbanSound8K [ 32 ]
Urban
✗
✓
✗
✗
AudioSet [ 12 ]
General
Partial
✓
✗
✗
Chen et al. [ 5 ]
Bathroom
✓
✗
✗
✗
Hyun [ 15 ]
Water/bathroom
✓
✗
✗
✗
Öztürk et al. [ 26 ]
Restroom
✓
✓
✗
✗
Table 1 : Representative datasets for environmental sound recognition and bathroom monitoring. SnaanGhar7 adds mobility-aid events, an explicit non-target class, and raw-waveform edge baselines.
Figure 1 : Assistive-living use scenario of SnaanGhar7 . The selected sound classes represent bathroom activities relevant to an older adult living independently. Door and walker/crutch sounds provide entry, exit, and mobility cues, while tap, shower, and flush-related sounds correspond to common hygiene activities.
Property
Value
Classes
7
Recording environments
5 (residential, hostel, institutional)
Contributors
3
Original duration
≈ 385 min
Total segmented windows
21,387
Clip duration
1.51125 s (30,225 samples)
Table 2 : Summary of the SnaanGhar7 dataset and its embedded recording platform.
Class
Description
Acoustic characteristics
Flush
Toilet flushing
Broadband water flow with transient onset
Shower
Shower water flow
Continuous broadband noise
Bathroom Tap
Bathroom tap running
Continuous water flow, varying intensity
Basin Tap
Basin tap running
Localized continuous water flow
Door
Opening/closing/knocking
Short impulsive sounds
Walker/Crutch
Mobility-aid movement
Repeated transient impacts
Table 3 : Class definitions and dominant acoustic characteristics.
Figure 2 : Embedded recording platform used for data acquisition.
ID
Location
Type
E01
NIT Durgapur
Institutional
E02
Residential Home, Howrah
Residential
E03
Hall-6, NIT Durgapur
Hostel
E04
DS-1B, NIT Durgapur
Residential
E05
Hall-9, NIT Durgapur
Hostel
Table 4 : Recording environments.
Figure 3 : Overview of the SnaanGhar7 data collection, annotation, preprocessing, and benchmark-generation pipeline.
Figure 4 : Class distribution of the SnaanGhar7 dataset.
Figure 5 : Signal-level statistics per class (violin plots). Impulsive classes ( Door , Walker/Crutch ) show markedly higher kurtosis and zero-crossing rate than the sustained water classes, providing a purely time-domain axis of separation that a raw-waveform model can exploit.
Figure 6 : Per-class mean power spectra. The sustained water classes ( Flush , Shower , Bathroom Tap ) overlap in the mid-frequency band, foreshadowing the model’s principal confusions, whereas impulsive classes carry distinct high-frequency transient energy.
Figure 7 : Representative mel-spectrogram per class. Impulsive events ( Door , Walker/Crutch ) appear as brief vertical onset stripes; the sustained water classes fill the window with visually similar broadband energy, previewing their confusability.
Figure 8 : Per-class temporal energy profile. Impulsive events ( Door , Walker/Crutch ) show a sharp onset and decay; sustained water events maintain a flat envelope. The onset structure gives a time-domain cue that complements the overlapping spectra of Fig. 6 .
Figure 9 : Per-class spectral descriptors (centroid, bandwidth, roll-off, flatness). Impulsive classes show higher centroids and flatness; sustained water classes cluster at lower centroids with smoother spectra, separating the two families along the same axis the model exploits.
Figure 10 : Within-class variability per class. The sustained water classes are internally variable (flow rate, basin geometry) as well as mutually overlapping, whereas impulsive classes are tighter. The combination of high intra-class spread and high inter-class overlap that makes the water triad the hardest group.
Class
Dc
Mean overlap
Intra spread
Flush
0.74
0.50
0.92
Door
0.54
0.42
0.79
Basin Tap
0.45
0.46
0.81
Shower
0.43
0.47
0.83
Walker/Crutch
0.31
0.38
0.77
Bathroom Tap
0.25
0.42
0.76
Table 5 : Per-class difficulty ranking from the signal-level analysis (higher = predicted harder), computed before model training. The aggregate score is a heuristic: its class-level correlation with model error is weak ( ρ=−0.07 ), while pairwise overlap is more useful for explaining specific confusions (Section 7.7 ).
Figure 11 : Per-class difficulty profiles across the five component diagnostics of Eq. ( 2 ) (inter-class overlap, intra-class spread, class imbalance, silence fraction, energy variability). Each class is hard for a different combination of reasons; Flush scores high on both overlap and spread, while the impulsive classes are driven by silence and energy variability.
Figure 12 : Linear-discriminant projection of the hand-crafted feature space. Impulsive classes and Basin Tap separate cleanly, while the sustained water classes overlap even under the best linear projection, an intrinsic, model-independent property of the acoustics.
Figure 13 : SincDPNet architecture. A learnable sinc filter bank, parameterized by per-filter cutoff and bandwidth in Hz, maps the raw waveform to Nf band-pass channels. A depthwise-separable convolutional body then performs classification. The front end contains only 2Nf trainable parameters.
Component
Parameters
Share
Sinc front-end ( 2Nf )
50
1.76%
Depthwise-separable body + head
2,798
98.24%
Total
2,848
100%
Table 6 : Parameter budget of the compact SincDPNet ( Nf=25 ). The filter-bank parameter count is independent of kernel length L .
Figure 14 : Quality-based coreset construction. (a) Distribution of the composite quality score across training clips; the top 25% per class is retained for search. (b) Validation macro-F1 for the same candidate architectures under coreset and full-data training. The reported rank correlation indicates how well the coreset preserves the ordering used during optimization.
Variable
Symbol
Range
Primarily affects
Sinc filters
Nf
16 – 128
size, F1, resolution
Sinc kernel length
L
101 – 401 (odd)
latency, F1
Body width mult.
α
2 – 12
size (dominant)
DS-conv blocks
B
3 – 6
size, latency
Peak LR
η
10−3 – 3×10−2
F1 (convergence)
Pooling stride
s
{2,4,8}
latency, F1
Table 7 : Design variables for the MOBO case study.
Figure 15 : Hypervolume versus evaluation for Study A. The surrogate-guided configurations reach similar final hypervolume values and outperform random search within the fixed evaluation budget.
Config
Acquisition
Kernel
Final HV
Evals to 90%
C1
qEHVI
Matérn-5/2
2091.27
17
C2
qNEHVI
Matérn-5/2
2090.65
18
C3
qParEGO
Matérn-5/2
2089.43
22
C4
qEHVI
RBF
2090.12
19
C5
Random
—
2078.49
>24
Table 8 : Acquisition and kernel sensitivity for Study A at a fixed budget of 24 model evaluations. Higher final hypervolume (HV) is better, while “Evals to 90%” indicates the evaluations required to reach 90% of the best HV gain. Best values are shown in bold .
Figure 16 : Pareto fronts from the two design studies. Dominated configurations are shown faded; points are colored by weighted score, and the highest-F1 and highest weighted-score selections are marked.
Rank
Architecture
Performance
Selection
Nf
α
B
Size (KB)
Macro-F1
Weighted
Pareto
1
78
4
5
13.3
0.8443
0.8523
⋆
2
25
6
4
11.1
0.8016
0.7842
⋆
3
24
8
5
36.7
0.8592
0.7698
⋆
4
20
2
3
1.4
0.7622
0.7574
⋆
5
24
3
5
7.9
0.7783
0.7562
⋆
Table 9 : Study-A configurations ranked according to the weighted score. Nf denotes the number of sinc filters, α denotes the width multiplier, and B denotes the number of depthwise-separable blocks. A ⋆ indicates a Pareto-optimal configuration in the Macro-F1–model-size plane. Bold values indicate the best overall values.
Model
Input
Params
Acc. (%)
Bal. Acc.
Macro-F1
MCC
DS-CNN [ 36 ]
MFCC
18,311
75.3 ± 6.8
0.811 ± 0.053
0.705 ± 0.063
0.715 ± 0.071
CRNN [ 3 ]
MFCC
236,967
74.6 ± 4.3
0.826 ± 0.024
0.695 ± 0.027
0.709 ± 0.041
TinyCNN [ 36 ]
MFCC
93,575
76.5 ± 7.4
0.809 ± 0.058
0.698 ± 0.070
0.729 ± 0.076
SincNet [ 30 ]
Raw
132,327
76.7 ± 1.8
0.809 ± 0.027
0.712 ± 0.013
0.721 ± 0.023
ACDNet [ 24 ]
Raw
22,055
71.7 ± 3.0
0.729 ± 0.032
0.639 ± 0.033
0.670 ± 0.028
DPNet [ 6 ]
Raw
4,944
69.3 ± 2.0
0.689 ± 0.016
0.612 ± 0.009
0.645 ± 0.020
Table 10 : Performance comparison on SnaanGhar7 under environment-disjoint splitting. Results are reported as mean ± standard deviation over three random seeds. Parameter counts are exact. The best result in each metric is shown in bold.
Figure 17 : Test-set t-SNE embeddings of the two configurations selected by the search. The best-F1 model (a) uses more capacity and pulls the water classes slightly further apart, while the compact rank-1 model (b) keeps the same overall layout at a fraction of the size. In both, the impulsive classes separate cleanly and the sustained water classes share one region, so the structure is preserved across the trade-off.
ID
Variant
Macro-F1
Δ vs. A0
A0
SincDPNet (reference: Nf=25 , L=401 )
0.651 ± 0.049
0.000
A1a
Fixed (non-learnable) sinc front-end
0.644 ± 0.046
−0.007
A1b
Plain trainable conv front-end
0.615 ± 0.051
−0.036
A1c
MFCC + same body
0.636 ± 0.044
−0.015
A2a
Nf=16
0.613 ± 0.052
−0.038
A2b
Nf=32
0.628 ± 0.047
−0.023
Table 11 : Component ablations under environment-disjoint splitting (macro-F1, mean ± standard deviation over three seeds). Δ denotes the change in mean macro-F1 relative to A0. Best macro-F1 is shown in bold .
Figure 18 : Learned SincDPNet filter bank (each bar is one band-pass filter; width = bandwidth). The learned allocation of spectral resolution is directly readable in Hz.
Figure 19 : Overlaid magnitude responses of the full learned filter bank on a shared frequency axis. The dense overlap in the low-to-mid band and the sparse coverage of the high band show, in a single view, how the bank concentrates spectral resolution where the water classes carry their energy.
Figure 20 : Frequency-magnitude responses of the learned sinc filters. The band-pass channels tile the low-to-mid frequency range densely and the high range sparsely, the frequency-domain view of the allocation in Fig. 18 .
Figure 21 : Per-class mean activation energy across learned filters. Sustained water classes share active bands (explaining their confusability), whereas Basin Tap and impulsive classes occupy distinct bands.
Figure 22 : Per-class activation traced across individual learned filters. The sustained water classes activate nearly the same filters, while the impulsive classes and Basin Tap each drive a separable group, mirroring the confusion structure of Fig. 24 .
Figure 23 : Acoustic-to-model comparison for the compact Nf=25 model. Left: the aggregate class-difficulty score has little association with class error. Right: pairwise acoustic overlap has a modest positive association with off-diagonal confusion.
Figure 24 : Row-normalized (recall, %) confusion matrices of the two configurations selected by the multi-objective search. The best-F1 model (a) recovers Basin Tap and Door , the classes the compact model sacrifices, at the cost of 12× more footprint; the rank-1 weighted model (b) preserves the easy classes at 13.3 KB but leaves Basin Tap unresolved. Added capacity is spent precisely on the hardest, most overlapping classes.
Figure 25 : Pairwise Bhattacharyya overlap (Eq. 1 ) between class feature distributions ( 1= identical). The sustained water classes form the highest-overlap block, providing the a priori prediction that the trained model later confirms in its confusion matrix (Fig. 24 ).
Figure 26 : t-SNE embeddings of the held-out test set for all benchmarked models. The sustained water classes form a single overlapping region in every model regardless of size, while the impulsive classes and Basin Tap stay separated, evidence that the confusability structure is a property of the data, not of any one architecture.
Selection role
Nf
α
Blocks
Val. macro-F1
Size (KB)
Highest weighted score (rank-1)
78
4
5
0.8443
13.3
Highest validation F1
59
10
5
0.8942
54.8
Compact Pareto model
25
6
4
0.8016
11.1
Table 12 : Study-A configurations selected for deployment-aware evaluation. The three configurations represent distinct points on the Pareto frontier: best weighted score (rank-1), highest validation macro-F1, and a compact model.
Device (RAM, clock)
Model
Params
Acc. (%)
Macro-F1
Inf. time
RTF
Raspberry Pi Zero 2 W ( 512 MB, 1 GHz)
DPNet (baseline)
4,944
69.2
0.612
15 ms
0.01
SincDPNet, compact ( Nf=25 )
2,848
75.7
0.661
435.6 ms
0.29
SincDPNet, best-F1
14,040
80.2
0.760
920.4 ms
0.61
SincDPNet, rank-1 (weighted)
3,408
75.6
0.673
886.3 ms
0.59
Laptop (Intel Core U5) ( 16 GB, 1.2 GHz)
DPNet (baseline)
4,944
69.2
0.612
2.5 ms
0.0016
SincDPNet, compact ( Nf=25 )
2,848
75.7
0.661
28.0 ms
0.0186
Table 13: Single-clip inference performance of DPNet and selected SincDPNet configurations on embedded and reference hardware for a 1.5 s audio clip. Inference time excludes audio capture; RTF <1 indicates faster-than-real-time operation.
Error category
Count
Transient/impulsive overlap
774
Broadband water-flow ambiguity
406
Non-target/unknown confusion
29
Low SNR / background masking
0
Annotation-boundary (partial event capture)
0
Total
1,209
Table 14: Error characterization of the deployed SincDPNet over its T=1,209 misclassified test clips, assigned to mutually exclusive categories by an automated rule-based audit (each clip is counted once, under the first matching rule). Two acoustic categories account for 97.6% of all errors.
Passive Acoustic Monitoring of bats generates massive ultrasonic datasets (>27 GB/night per node), straining edge storage and battery life. Legacy triggers fail against acoustic confusers, while deep models exceed microcontroller limits. We present a hardware-aware Bat Activity Detector (BAD) specifically designed to discriminate bat calls from hard biological and environmental confusers across variable sampling rates (192-384 kHz). Tailored for the Silicon Labs EFM32PG26 (MVP) in 8-bit integer precision, our model achieves 100 percent hardware offload across all 14 layers (17.2 KB Flash, 73.1 KB RAM). End-to-end preprocessing (74.00 ms for 76 frames) and inference (30.00 ms) of 100 ms clips at 192 kHz require 104.00 ms per clip. On spatially out-of-domain recordings under a realistic low-prevalence regime (r_pos = 0.05), BAD achieves an AUC-ROC of 0.9748 and suppresses 99.4% of non-target noise frames while retaining 65.3% of bat calls - delivering a >33x precision gain over classical Goertzel baselines.
Stefano Ciapponi, Santiago Martinez Balvanera, Andrea Cesaretti +2
Fondazione Bruno Kessler · University of Trento · University College London
Passive Acoustic Monitoring (PAM) is an efficient and non-invasive method for surveying ecosystems at a reduced cost. Typically, autonomous recorders allow the acquisition of vast bioacoustic datasets which are then analyzed. However, power consumption and data storage are both scarce and limit the duration of acquisition campaigns. To address this issue, we propose a smart PAM system which allows the in-situ analysis of the soundscape by embedding a classifier directly onto an AudioMoth microcontroller. Specifically, we propose an optimized yet simple 1D Convolutional Neural Network (1D-CNN) to classify the raw audio. The model focuses on the specific call of Scopoli Shearwater seabirds (endangered species) and is trained on a real-world dataset with a classification accuracy of 91% (balanced accuracy of 89%). We also propose a process to optimize the model to fit the severe resource constraints of the AudioMoth, achieving a ~10kB RAM memory footprint and 20ms inference time. Finally, we present an open-source tutorial of our model optimization and export strategy which can be used for embedding models beyond the scope of our study. Our modified version of the AudioMoth firmware adds two functions: (F1) which selectively records data when the target species has been detected and (F2) which logs the continuous classification results in real time. This work intends to facilitate the conception of intelligent sensors, enhancing the efficiency and scalability of bioacoustic monitoring campaigns.
Louis Lerbourg, Paul Peyret, Juliette Linossier +1
Univ. Grenoble Alpes, CEA, List, Grenoble, France · Biophonia, France
We propose a compact underwater acoustic classification framework combining multi-representation feature engineering, temporal statistical pooling, and compact convolutional architectures designed for acoustic time-frequency and cochlear representations. We investigate multiple conventional and auditory-inspired representations and first evaluate lightweight classifiers and Conventional Neural Networks (CNNs) on ShipsEar dataset. On the provided split, a two-layer CNN achieves a macro F1 of 0.9918, while a Radial Basis Function Support Vector Machine (RBF-SVM) reaches 0.9883. However, source-recording provenance cannot be reconstructed, preventing verification of recording-independent generalisation. We therefore evaluate on DeepShip dataset using recording-level partitioning before segmentation. Under this protocol, a 157K-parameter compact CNN achieves a test macro F1 of 0.7226, while an 11.17M-parameter ResNet18 provides no improvement in validation performance under the matched setting. These results demonstrate the importance of representation-aware feature and model design, together with rigorous recording-level evaluation, for classification performance and deployability in compact underwater acoustic systems.
Abishek Soti, Thura Pyae Sone, Naqib Ibnul +4
ICNS, Western Sydney University, Sydney, Australia