People with sound sensitivity (PWSS) often manage distressing sounds with earplugs and noise-canceling headphones that broadly suppress their surroundings, limiting access to useful auditory cues. We present Sona, a mobile system for personalized, real-time soundscape mediation, informed by prior sound sensitivity research and an online survey of 68 PWSS. Sona selectively attenuates multiple overlapping user-chosen sounds at adjustable strength, suggests targets from ambient sound recognition, and lets users add custom targets from short recordings without retraining the model. In an in-situ evaluation with ten PWSS, participants reported that Sona made their soundscapes more manageable. The study also surfaced uneven attenuation across sound type and context, tensions between managing filters and attending to ongoing activities, and difficulty interpreting personalization outcomes. These findings highlight the need to design for the quality of the residual soundscape, balance user control with interaction demands, and support guided, interpretable personalization.
Figures & tables
Figure 1. Sona helps people with sound sensitivity manage challenging everyday soundscapes. (A) In a shared home, overlapping sounds such as keyboard typing, air conditioner hum, blender noise, and music can make it difficult to focus. (B) Conventional noise-canceling headphones broadly suppress the soundscape, reducing both bothersome and desired sounds. (C) Sona instead selectively attenuates bothersome sounds. Users can also personalize Sona by adding new sound targets from recordings of their environments. A three-part overview figure illustrating Sona, a system for helping people with noise sensitivity manage everyday acoustic environments. Panel A shows a user at a table in a shared home setting with a laptop user in the background, a blender in the kitchen, an air conditioner on the wall, and a speaker playing music. The panel text says the user struggles to focus in a noisy home environment, with environmental sounds including keyboard, air conditioner, blender, and music. Panel B shows conventional noise-canceling headphones treating all sounds the same: keyboard, air conditioner, blender, music, and speech are all muted, with no distinction between bothersome and desired sounds. Panel C shows Sona selectively attenuating bothersome sounds while preserving desired sounds. Keyboard, air conditioner, and blender are shown as bothersome sounds, while music and speech are preserved as desired sounds. A smaller subpanel below shows personalization over time: an unrecognized sound can become a new sound class, labeled “Blender,” allowing the system to be extended with sounds from the user’s own environment.
Modification Preference
Percentage
Lower / reduce the sound
36.2%
Completely remove the sound
29.1%
Change the tone and style
16.5%
Play pleasant sounds on top
15.3%
Leave them as is
13.7%
Table 1. Participant preferences for modifying unwanted sounds (multi-selection allowed).
Figure 2. Distribution of participants’ reported bothersomeness across sound categories (Top 5, sorted by ratings of very bothersome or above). A horizontal stacked bar chart showing the distribution of bothersomeness ratings across five sound categories. Each bar is divided into five segments: Not at all bothersome (dark blue), Slightly bothersome (light blue), Moderately bothersome (yellow), Very bothersome (orange), and Extremely bothersome (dark orange). Loud sounds: 8% not at all, 23% slightly, 14% moderately, 38% very, 17% extremely. Sharp sounds: 9% not at all, 23% slightly, 20% moderately, 22% very, 26% extremely. Tools: 35% slightly, 18% moderately, 16% very, 26% extremely. Mouth sounds: 18% not at all, 26% slightly, 14% moderately, 21% very, 21% extremely. Appliance sounds: 12% not at all, 25% slightly, 25% moderately, 33% very. Loud and Sharp categories show the highest combined proportions of very and extremely bothersome ratings.
Figure 3. Sona’s User Interface. (A) In Live Listen, Sona surfaces suggestions for detected sounds, lets users select target sounds, and supports real-time adjustment of attenuation strength. (B) When an unfamiliar sound matches a stored sensitivity profile, Sona prompts the user to save it for later personalization. (C) Users can manage custom sound classes and sensitivity profiles. (D) Users can create a reusable custom class by recording audio samples. Four iPhone screenshots showing Sona's user interface. Panel A shows the Live Listen view with a system suggestion banner reading "Dog barking detected, Tap to attenuate this sound," two active target sounds (Car alarms and Keyboard typing) each with a remove button, a Filter Strength slider set to 100%, and a Sound Recognition section showing "Dog" detected at 51% confidence. Panel B shows the Live Listen view with a Characteristic Match alert warning that a sound matches the user's "Whirring" sensitivity profile, with Save and Dismiss buttons. One target sound (Dog barking) is active, Filter Strength is at 100%, and Sound Recognition shows "Click" at 55% confidence. Panel C shows the Personalization view with two sections: Custom Classes listing "Car alarms" (3 samples, Ready) and "Blender" (2 samples, Ready) with an Add Custom Class button, and Personal Noise Sensitivity Profiles listing "Whirring" described as "High-pitched, whirring sounds from appliances" and "Squeaking sounds" described as "Any type of chair or shoe squeaking," with an Add Characteristic button. A Memory section below shows two saved audio clips labeled "Whirring" with natural-language descriptions and timestamps. Panel D shows the Custom Class creation view for "Car alarms" with three recorded audio samples (6.4s, 4.3s, and 3.8s), a recording interface showing 1.6 of 10 seconds elapsed with a red stop button, and an Embedding section showing a green checkmark indicating the embedding is ready for use.
Attenuation Quality
Targets
Input SI-SNR
Output SI-SNR
SI-SNRi
1
5.01
8.30
3.29
2
1.54
4.54
3.00
3
−0.37
2.86
3.23
Table 2. Attenuation benchmark results: mean SI-SNR (dB) between the retained-scene reference and the unprocessed mixture (Input), the processed output (Output), and their difference (SI-SNRi), for one to three simultaneous targets.
Metric
H1
H2
H3
P1
P2
P3
C1
C2
C3
CPU Usage (%)
17.88
17.76
18.08
18.69
23.08
23.05
17.34
18.08
17.55
Inference Time (ms)
11.32
11.30
11.34
10.97
10.98
11.15
11.22
11.15
11.21
E2E Latency (ms)
42.03
42.02
42.01
41.56
41.61
41.70
41.74
41.56
41.72
Battery Usage (%/hr)
16.60
16.70
17.20
17.70
18.60
17.00
16.50
16.50
16.60
Table 3. Runtime performance of Sona across nine consecutive 10-minute runs. H, P, and C denote the home, public indoor, and cityscape scenes described in Section 4.1.5 ; the numeral gives the number of simultaneous targets (1–3).
Figure 4. Sona’s two-layer system architecture. The Live Mediation Layer runs on-device: users select target sounds manually or accept suggestions from a sound recognizer, and the neural attenuation model reduces the selected sounds at an adjustable strength while preserving the rest of the scene. The Personalization Layer extends the built-in sound set: users describe personal triggers as sensitivity profiles, Sona matches audio-language-model descriptions of the live soundscape against these profiles and prompts users to save matching clips, and users create custom sound classes from saved or manually recorded examples. Two-layer diagram beside a kitchen photo in which a barking dog is labeled "Dog Barking" and a running blender is labeled "Blender." In the Live Mediation Layer, a Sound Recognizer box points to a Trigger Suggestion box with the example prompt "Dog Barking detected. Suppress it?" Both the Trigger Suggestion and Manual Selection boxes point to an On-Device Neural Attenuation box. In the Personalization Layer, a Personal Sensitivity Profiles box points to a Sensitivity Profile Matching with Audio LM box with the example prompt "I hear mechanical whirring sounds, which match your sensitivity profiles. Save it?"; it points down to Discovered Trigger Clips, which points to Custom Sound Class.
Figure 5. Sona’s trigger discovery pipeline. When an unsupported sound is detected, the system generates a description and matches it against user-defined sensitivity profiles to determine whether it should be saved for later personalization. Diagram of Sona's unsupported-sound discovery pipeline. An environmental sound is analyzed by Apple's SoundAnalysis and by Audio Flamingo 3, which produces a rich audio description. The classification result and description are compared against the user's stored sensitivity profiles by an on-device LLM judge. If a profile matches, the interface presents a lightweight save-or-dismiss suggestion.
Table 4. Participant demographics, SSSQ-2 symptom profiles, and current coping strategies. Profiles list forms of sound sensitivity with scores above zero; parenthetical values show the obtained score and maximum possible score. Hyperacusis (0–12) encompasses loudness intolerance, sound-induced pain or discomfort, and fear of worsening hearing or tinnitus. Misophonia (0–3) reflects anger or anxiety elicited by specific trigger sounds, while noise sensitivity (0–3) reflects disturbance from environmental noise. Scores represent self-reported symptom frequency during the preceding two weeks rather than clinical diagnoses. A score of zero, corresponding to the 0–1 days response category, is omitted.
Figure 6. Participant ratings ( n=10 ) on 7-point agreement scales (1 = strongly disagree, 7 = strongly agree). Stacked bars show response counts, with means ( M ) and standard deviations ( SD ) on the right. Top : Perceived usefulness of selective suppression and ease of using the interface, rated after each in-situ scene, and usefulness and ease of use of the Personalization Layer, rated after the personalization activity. Each measure used a single item adapted from UMUX-Lite ( Lewis et al., 2013 ) . Bottom : Experience ratings collected after each scene for three statements, shown in shortened form in the figure: “Sona reduced the sounds that were bothering me,” “I could still hear the sounds I wanted or needed to hear,” and “Sona made the sound environment more comfortable.” Two panels of horizontal stacked bars show ratings from 10 participants on a seven-point scale from strongly disagree to strongly agree. Segments show response counts, with means and standard deviations beside each bar. The top panel shows perceived usefulness and ease of use. Mean usefulness ratings for selective suppression are 5.30 in the lounge and 4.90 in the food court; interface ease-of-use ratings are 6.10 and 5.50, respectively. Personalization receives mean ratings of 5.60 for usefulness and 5.90 for ease of use. The bottom panel shows three experience ratings for each setting. Mean ratings in the lounge and food court, respectively, are 5.90 and 5.30 for reducing bothersome sounds, 5.50 and 4.70 for hearing wanted or needed sounds, and 5.80 and 5.60 for making the environment more comfortable. Responses generally favor agreement, with higher average ratings in the lounge across all scene-specific items. Hearing wanted or needed sounds in the food court receives the lowest average experience rating and the most neutral or disagreeing responses.
Misophonia, a disorder of decreased tolerance to specific sounds, affects 5-20% of the population, yet sufferers have no good options: therapy helps a minority, and earplugs or noise cancellation silence everything. We present a study for neural trigger sound suppression for misophonia, selectively removing trigger sounds. We curate a dataset covering the 10 most common trigger classes. Using streaming dual-path networks operating on 6 ms audio chunks, we explore both one-hot and multi-hot-conditioned models that suppress 1-3 triggers from the acoustic scene. We validate our model outputs in a listening study with 30 adults with clinically elevated misophonia impairment. Participants reported significantly lower distress and arousal, and improved valence, for suppressed audio.
Speaker anonymization (SA) systems modify timbre while leaving regional or non-native accent cues intact, which is problematic because such cues can reveal a speaker's first-language or geographic background and narrow the anonymity set. To address this issue, we present PHONOS, a streaming module for real-time SA that performs accent neutralization in a privacy sense: reducing accent-origin cues by converting non-native segmental realizations toward a chosen target accent domain. Our approach pre-generates golden speaker utterances that preserve source timbre and rhythm but replace foreign segmentals with native ones using silence-aware DTW alignment and zero-shot voice conversion. These utterances supervise a causal accent translator that maps non-native content tokens to native equivalents with at most 40ms look-ahead, trained using joint cross-entropy and CTC losses. Our evaluations show an 81% reduction in non-native accent confidence, with listening-test accentedness ratings consistent with this shift. PHONOS also moves outputs away from the original speaker in embedding space, suggesting lower linkability under an embedding-based proxy, while running with ≤241ms end-to-end latency on a single GPU.
Waris Quamer, Mu-Ruei Tseng, Ghady Nasrallah +1
Department of Computer Science & Engineering, Texas A&M University, College Station, US
We present an agentic soundscape construction framework for controllable compositional audio generation that makes explicit the scene planning, source selection, temporal layout, and rendering steps typically handled implicitly by single-shot text-to-audio models. An LLM-based agent converts user intent into an executable scene plan, acquires assets through retrieval and on-demand generation, renders controllable multi-event mixtures, and exports aligned scene metadata. The framework also supports human-in-the-loop interaction through user-guided tool selection and editable scene plans. Together, these components provide an inspectable and reusable approach to controllable soundscape synthesis and scalable audio-language data construction. Listener studies and objective metrics demonstrate competitive generation performance against text-to-audio baselines, while models trained with agent-generated data consistently outperform real-only baselines in downstream audio reasoning. Code, demos, and listening-test results are available at https://haozhang6720.github.io/SoundscapeAgentDemoPage/.
Hao Zhang, Yiwen Zhao, Yixuan Zhang +2
Wuhan University, Wuhan, China · Tencent Hunyuan, Bellevue, WA, USA