PLACE: Positional Latent Adaptation via Conditioned Embeddings for Binaural Audio Generation
Authors: Tiernon Riesenmy, You Zhang, Gautam Bhattacharya, Andrea Fanelli
Organizations: The Wallace H. Coulter Department of Biomedical Engineering, Georgia Institute of Technology, Atlanta, GA, USA · Advanced Technology Group, Dolby Laboratories, San Francisco, CA, USA
We present PLACE, a method that extends the pretrained any-to-audio model AudioX for binaural generation from arbitrary combinations of text, video, and optional audio prompts. PLACE augments video conditioning with Perception Encoder Core features, aligns text and video representations to derive spatial cues, and applies a conditioning-dependent low-rank transformation to the generated latent. The adapter is supervised via decoded-audio interaural level and time difference objectives. Trained on MRSAudio, PLACE improves most metrics over ViSAGe on FAIR-Play and achieves an improved SpatialCLAP score over SpatialSonic on the BEWO-1M Single Static test split. Listener evaluations favor PLACE for video-to-audio and out-of-distribution text-to-audio generation, demonstrating flexible multimodal control and improved spatial consistency.
Figures & tables
Figure 1: Overview of our proposed PLACE framework. We emphasize the video feature path in blue and the text feature path in orange. a. AudioX generates a latent; spatial conditioning drives its transformation before the frozen SAO decoder. All five pretrained encoders and the decoder remain frozen; AudioX fine-tuned via LoRA. b. Projected PE Core features augment CLIP and Synchformer features before the MAF module. c. Text–video cross-attention and joint self-attention yield spatial features Hc-sp .
FAIR-Play V2A
Split / Model
FSAD
ITD
ILD
SCLAP
1 / ViSAGe
15.9016
0.1136
2.1306
–
1 / PLACE
15.2879
0.1077
2.0495
–
2 / ViSAGe
14.8996
0.1133
2.0198
–
2 / PLACE
14.3745
0.1217
1.9827
–
3 / ViSAGe
14.2590
0.1228
2.1849
–
Table 1: Out-of-distribution objective evaluation on FAIR-Play (video-to-audio, three test splits) and the BEWO-1M SS-set test split (text-to-audio). Lower is better for all metrics; ITD and ILD are reported in ms and dB; SCLAP is the SpatialCLAP difference from ground truth. Bold marks the better score between PLACE and the baseline in each row.
Model
FSAD
ITD
ILD
SCLAP
PLACE
5.14
0.11
1.98
0.04
LoRA + PE
6.30
0.11
2.17
0.18
LoRA only †
7.10
0.12
2.25
0.15
LoRA + PE + I
7.22
0.12
2.36
0.21
LoRA + PE + I/T + MD
8.96
0.13
2.58
0.23
LoRA + PE †
9.20
0.12
2.20
0.20
Table 2: Ablation study on the held-out MRSAudio set ( n=596 ). Lower is better for all metrics. PE: PE Core video features; I/T: ILD/ITD losses; MD: modality dropout; † : audio and video conditioning only (no text). Bold marks the best score in each column.
Model
SC
SI
SP
Sy
AR
Avg
T2A / In-distribution ( n=24 )
GT
54.0
65.8
52.1
–
69.6
60.4
PLACE
38.3
38.6
31.3
–
53.8
40.5
SpatialSonic
44.2
48.2
42.5
–
42.7
44.4
AudioX
49.2
33.0
27.8
–
54.4
41.1
T2A / Out-of-distribution ( n=25 )
Table 3: Subjective listener evaluation. Ratings range from 1–100, higher is better. SC: semantic consistency; SI: spatial impression; SP: spatial consistency; Sy: synchronization (omitted for T2A); AR: audio realism; Avg: overall average across the rated dimensions. GT: ground truth, shown as a reference and excluded from bolding. Bold marks the best model score in each column within each condition.