Relational Synthesis: Structure-Mediated Concatenative Synthesis for Foley and Retrieval-Augmented Audio Generation
Organizations: University of California San Diego, La Jolla, CA, USA
Abstract
We ask: given a retrieved source audio and a separate reference audio , can we synthesize novel audio out of this pair such that remains acoustically consistent with , while not persistently copying segments of or ? The first clause is a well-known goal in Foley audio production, and the second is a well-known issue in neural RAG when and are naively injected into neural generators. We show that both clauses can be addressed simultaneously using a method we coin relational synthesis, a variation of concatenative synthesis where target cost is replaced by a relational Gromov-like structural cost. Rather than imitating the content of , relational synthesis exploits it from the "other side of the hill": it transfers the temporal structure and directed amplitude motion of to reorganize and concatenate the grains of in a novel manner that protects 's acoustic information. Our experiments show that relational synthesis integrates naturally with neural RAG and produces Foley audio that performs well on metrics measuring temporal agreement, acoustic fidelity, and leakage persistence, while maintaining distribution-level quality and text alignment.
Figures & tables
| Temporal | Acoustic | Leakage | Prompt | Distribution | ||||||||||
| Condition | ||||||||||||||
| None | .086 | .094 | .354 | .305 | .913 | .540 | 1.806 | .542 | .167 | .177 | .118 | .246 | 8.83 | 8.82/.413/52.04 |
| Direct reference-audio controls | ||||||||||||||
| Ref | .083 | .059 | .402 | .421 | .939 | .524 | 2.196 | .511 | .170 | .243 | .365 † | .227 | 7.97 | 11.24/.405/56.74 |
| Mix | .080 | .084 | .395 | .347 | .927 | .562 | 1.727 | .590 | .303 | .198 | .196 † | .240 | 9.21 | 7.49/ .354 /47.31 |
| Source-carrier conditions – central comparison | ||||||||||||||