PLACE: Positional Latent Adaptation via Conditioned Embeddings for Binaural Audio Generation
Authors: Tiernon Riesenmy, You Zhang, Gautam Bhattacharya, Andrea Fanelli
Organizations: The Wallace H. Coulter Department of Biomedical Engineering, Georgia Institute of Technology, Atlanta, GA, USA · Advanced Technology Group, Dolby Laboratories, San Francisco, CA, USA
We present PLACE, a method that extends the pretrained any-to-audio model AudioX for binaural generation from arbitrary combinations of text, video, and optional audio prompts. PLACE augments video conditioning with Perception Encoder Core features, aligns text and video representations to derive spatial cues, and applies a conditioning-dependent low-rank transformation to the generated latent. The adapter is supervised via decoded-audio interaural level and time difference objectives. Trained on MRSAudio, PLACE improves most metrics over ViSAGe on FAIR-Play and achieves an improved SpatialCLAP score over SpatialSonic on the BEWO-1M Single Static test split. Listener evaluations favor PLACE for video-to-audio and out-of-distribution text-to-audio generation, demonstrating flexible multimodal control and improved spatial consistency.
Figures & tables
Figure 1: Overview of our proposed PLACE framework. We emphasize the video feature path in blue and the text feature path in orange. a. AudioX generates a latent; spatial conditioning drives its transformation before the frozen SAO decoder. All five pretrained encoders and the decoder remain frozen; AudioX fine-tuned via LoRA. b. Projected PE Core features augment CLIP and Synchformer features before the MAF module. c. Text–video cross-attention and joint self-attention yield spatial features Hc-sp .
FAIR-Play V2A
Split / Model
FSAD
ITD
ILD
SCLAP
1 / ViSAGe
15.9016
0.1136
2.1306
–
1 / PLACE
15.2879
0.1077
2.0495
–
2 / ViSAGe
14.8996
0.1133
2.0198
–
2 / PLACE
14.3745
0.1217
1.9827
–
3 / ViSAGe
14.2590
0.1228
2.1849
–
Table 1: Out-of-distribution objective evaluation on FAIR-Play (video-to-audio, three test splits) and the BEWO-1M SS-set test split (text-to-audio). Lower is better for all metrics; ITD and ILD are reported in ms and dB; SCLAP is the SpatialCLAP difference from ground truth. Bold marks the better score between PLACE and the baseline in each row.
Model
FSAD
ITD
ILD
SCLAP
PLACE
5.14
0.11
1.98
0.04
LoRA + PE
6.30
0.11
2.17
0.18
LoRA only †
7.10
0.12
2.25
0.15
LoRA + PE + I
7.22
0.12
2.36
0.21
LoRA + PE + I/T + MD
8.96
0.13
2.58
0.23
LoRA + PE †
9.20
0.12
2.20
0.20
Table 2: Ablation study on the held-out MRSAudio set ( n=596 ). Lower is better for all metrics. PE: PE Core video features; I/T: ILD/ITD losses; MD: modality dropout; † : audio and video conditioning only (no text). Bold marks the best score in each column.
Model
SC
SI
SP
Sy
AR
Avg
T2A / In-distribution ( n=24 )
GT
54.0
65.8
52.1
–
69.6
60.4
PLACE
38.3
38.6
31.3
–
53.8
40.5
SpatialSonic
44.2
48.2
42.5
–
42.7
44.4
AudioX
49.2
33.0
27.8
–
54.4
41.1
T2A / Out-of-distribution ( n=25 )
Table 3: Subjective listener evaluation. Ratings range from 1–100, higher is better. SC: semantic consistency; SI: spatial impression; SP: spatial consistency; Sy: synchronization (omitted for T2A); AR: audio realism; Avg: overall average across the rated dimensions. GT: ground truth, shown as a reference and excluded from bolding. Bold marks the best model score in each column within each condition.
Recent joint audio-video generative models can synthesize realistic videos with synchronized sound, but typically generate audio as a single mixed track. This limits source-level control and differs from practical audiovisual workflows, where speech, music, sound effects, and ambient sounds are represented as separate editable tracks. We introduce Soundwich, a training-free framework that transforms a frozen joint audio-video flow-matching model into a generator of multiple synchronized, independently editable audio stems coupled to a shared video. Soundwich generates separate audio stems with explicit control over their temporal activity. To keep separately generated sounds coherent, we introduce a shared scene representation that communicates global audiovisual context across stems while preserving their source-level separation. We further route cross-modal interactions between each audio stem and its corresponding visual source, improving audiovisual consistency. The resulting stems remain synchronized with the video and can be independently retimed, muted, replaced, or remixed. Experiments and human evaluations show improved temporal control, source separation, and naturalness, while enabling flexible source-level editing within coherent audiovisual generation. Code is available at https://github.com/CodyNing/Soundwich.
We propose a step-by-step video-to-audio (V2A) generation method that provides finer control over the generation process and more realistic audio synthesis. Inspired by traditional Foley workflows, our approach enables incremental generation of complementary sounds, allowing users to author multiple sound events induced by a video. To avoid the need for costly multi-reference video-audio datasets, each generation step is formulated as a negatively guided V2A process that discourages duplication of sounds already present in previously generated tracks. The guidance model is trained by finetuning a pre-trained V2A model on audio pairs from non-overlapping segments of the same video, encouraging it to leverage acoustic context while remaining visually grounded, and enabling training with standard single-reference audiovisual datasets. Objective and subjective evaluations demonstrate that our method enhances the separability of generated sounds at each step and improves the overall quality of the final composite audio, outperforming existing baselines. Our project page is available at: https://ahykw.github.io/sbsv2a/.
Akio Hayakawa, Masato Ishii, Takashi Shibuya +1
Sony AI, Tokyo, Japan · Sony Group Corporation, Tokyo, Japan
Audio generation has long been fragmented, with speech, music, and sound effects produced by domain-specific models that fail to jointly generate coherent audio scenes from a single description. The key obstacles are insufficient fine-grained supervision for real-world mixed audio and limited acoustic representations for modeling concurrent audio components. We present Dasheng AudioGen, a unified framework for generating general mixed-audio scenes from text. Dasheng AudioGen introduces structured multi-view captions, which explicitly decouple complex acoustic scenes into complementary description views, thereby enabling fine-grained control over audio layers. Furthermore, we employ a high-dimensional unified semantic-acoustic representation as the shared latent space. It injects semantic priors that facilitate cross-modal training convergence, while its high-dimensional feature space provides sufficient capacity to disentangle and fuse concurrent audio components effectively. With these designs, a simple flow-matching DiT achieves high-quality end-to-end audio scene generation. We also establish a comprehensive evaluation pipeline for audio scene generation. Experiments demonstrate that Dasheng AudioGen achieves performance approaching real-world recordings in mixed-audio categories, while remaining competitive with specialized models in single-type generation tasks. Demos are available at https://nieeim.github.io/Dasheng-AudioGen-Web/.
Jiahao Mei, Heinrich Dinkel, Yadong Niu +7
1X-LANCE Lab, Shanghai Jiao Tong University, Shanghai, China · 2MiLM Plus, Xiaomi Inc., Beijing, China