We ask: given a retrieved source audio S and a separate reference audio R, can we synthesize novel audio Y out of this pair (S,R) such that Y remains acoustically consistent with S, while not persistently copying segments of S or R? The first clause is a well-known goal in Foley audio production, and the second is a well-known issue in neural RAG when S and R are naively injected into neural generators. We show that both clauses can be addressed simultaneously using a method we coin relational synthesis, a variation of concatenative synthesis where target cost is replaced by a relational Gromov-like structural cost. Rather than imitating the content of R, relational synthesis exploits it from the "other side of the hill": it transfers the temporal structure and directed amplitude motion of R to reorganize and concatenate the grains of S in a novel manner that protects S's acoustic information. Our experiments show that relational synthesis integrates naturally with neural RAG and produces Foley audio that performs well on metrics measuring temporal agreement, acoustic fidelity, and leakage persistence, while maintaining distribution-level quality and text alignment.
Figures & tables
Figure 1: Relational synthesis pipeline. Reference R specifies temporal relations while source S supplies selectable acoustic material. Method M reorganizes source grains into carrier XM , which conditions pretrained generator Gθ together with prompt P to produce neural output Y . In our experiments, R is a vocal imitation paired with target T of the same class as S .
Figure 2: Out-of-dataset example with shakuhachi source S and humpback-whale reference R . Raw source conditioning exhibits persistent source copying; pointwise concatenation reduces copying but poorly follows R ’s temporal envelope. FO relational synthesis instead transfers the reference pattern while reorganizing source grains. Red boxes mark corresponding envelope structure.
Temporal
Acoustic
Leakage
Prompt
Distribution
Condition
LT↓
LR↓
OT↑
OR↑
MT↑
CT↑
KLT↓
CS↓
RS2s↓
CR↓
RR2s↓
CP↑
IS↑
FDV/C/P↓
None
.086
.094
.354
.305
.913
.540
1.806
.542
.167
.177
.118
.246
8.83
8.82/.413/52.04
Direct reference-audio controls
Ref R
.083
.059
.402
.421
.939
.524
2.196
.511
.170
.243
.365 †
.227
7.97
11.24/.405/56.74
Mix S+R
.080
.084
.395
.347
.927
.562
1.727
.590
.303
.198
.196 †
.240
9.21
7.49/ .354 /47.31
Source-carrier conditions – central comparison
Table 1: Stable Audio 3 results on 248 evaluation triplets. Pair-level metrics first average the three neural seeds of each (S,R,T) triplet; KLT , IS , and FDV/C/P are set-level. Bold marks the best non-shortcut source-carrier result. Underlining marks a numerical winner that is principally a shortcut or control bound. † marks persistence of a directly exposed waveform. Arrows indicate the favorable direction for each evaluation role; Leakage metrics are diagnostic rather than minimization objectives, so numerical winners are not highlighted.
We present FoleyGenEx, a unified video-to-audio (VTA) framework integrating multi-modal control, frame-level temporal alignment, and fine-grained semantics, enabling synchronized, versatile audio synthesis for diverse tasks. Existing VTA methods either have multi-modal control but weak temporal alignment or strong alignment but lack reference audio conditioning and semantic precision. FoleyGenEx fills this gap via three core innovations: a conditional injection mechanism for audio-controlled VTA and Foley extension, a multi-modal dynamic masking strategy preserving training synchronization, and an adverb-based data augmentation algorithm leveraging signal processing and large language models to enhance textual supervision with nuanced semantics. Experiments on AudioCaps, VGGSound, and Greatest Hits demonstrate its competitive controllable VTA performance against existing methods. Demo samples are available at https://foleygenex.github.io/FoleyGenEx.
Shiyao Wang, Xijuan Zeng, Hui Wang +4
Academy for Advanced Interdisciplinary Studies, Nankai University, Tianjin, China · Kling Team, Kuaishou Technology, Beijing, China
Recent unified audio generation models can support diverse tasks across speech, sound effects, and music, but most of them still focus on isolated task-level synthesis. However, real video production often requires multiple components of a complete audio track to be generated jointly and consistently for the same video. We present Foley-Omni, a unified multimodal audio generation model that extends isolated task-level synthesis to complete video soundtrack generation by jointly modeling speech, sound effects, and music within a shared latent generation process. To support training and reproducible evaluation, we develop an audiovisual data curation pipeline and introduce V2ST-Bench, a benchmark for holistic video soundtrack generation evaluation. Experiments show that Foley-Omni achieves competitive performance with expert systems on individual synthesis tasks, while improving speech intelligibility, audiovisual consistency and perceptual quality for mixed soundtrack generation.
Ye Tao, Lupeng Liu, Xuenan Xu +6
School of Intelligence Science and Technology, Nanjing University · Video Rebirth · Shanghai Jiao Tong University +2
Brain-to-audio reconstruction is limited by \emph{prior domination}: when a pretrained generator is conditioned on a weak neural signal, it produces realistic but stimulus-inaccurate audio. We introduce RAG-Audio, which decodes fMRI into a semantic audio embedding, retrieves a matching real-audio exemplar, and initializes the frozen generator's sampling trajectory from that exemplar while retaining the decoded embedding as conditioning. On Brain2Music, RAG-Audio improves 10-way stimulus identification from 0.14--0.18 for direct generation, near the 0.10 chance level, to 0.40--0.43, comparable to retrieval. It also reduces Fréchet Audio Distance by roughly an order of magnitude, from 13.49 to 1.25 for AudioLDM. RAG-Audio approaches nearest-neighbor retrieval in identification while remaining generative; its higher FAD is expected because retrieval directly replays real audio. An autoregressive negative control, which lacks an initializable latent trajectory, shows no comparable gain, attributing the improvement to trajectory initialization. These results suggest that retrieval-guided initialization can mitigate prior domination in brain-to-audio generation.