Creating sound effects for a new game-character skin requires a distinct acoustic identity while preserving gameplay-event roles. The challenge is to complete a coherent set of related sounds whose required degrees of redesign differ. We formulate this task as completion conditioned on base-skin audio, completed target assets, and a textual design description. We develop a pipeline to collect, process, and align corresponding events across League of Legends skins. Building on Stable Audio 3's pretrained audio prior, we fine-tune a latent inpainting model to jointly complete missing events. A signed soft retention mask encodes available audio and an adjustable transformation hint for each missing event, specifying the requested balance between retention and redesign. Experiments on held-out skins show improved reconstruction over the evaluated general-purpose audio editors. Target-derived hints further improve paired similarity, with three-level hints retaining most of the benefit of continuous guidance.
Table 1: Test results with requested k=2 (256 sequences, 687 unknown windows). Prefix denotes observed target assets; oracle hints use unknown target audio. Best values per block are bold. SA3: Stable Audio 3; a2a: audio-to-audio; η : noise strength.
Hint setting
Mel ↓
CLAP ↑
Constant h=0.5
9.39
0.660
Shuffled percentile hints
9.76
0.652
Three-level oracle hints
8.97
0.670
Positive similarity markers
9.29
0.670
Continuous oracle hints
8.93
0.673
Table 2: Hint ablations on the full test set at k=2 . All variants except the positive-marker model share one checkpoint.
Ours (oracle)
Copy-base
k
Mel ↓
CLAP ↑
Mel ↓
CLAP ↑
0
9.34
0.661
12.00
0.668
1
9.10
0.668
12.37
0.663
2
8.75
0.671
12.65
0.659
3
8.57
0.675
13.15
0.657
4
9.01
0.666
13.61
0.649
Table 3: Prefix sweep on the same 111 test sequences with at least five audible events, using 25 steps and continuous oracle hints. Scored suffix events change with k .