cs.SDSep 30, 2026

Game Sound-Effect Completion with Event-Level Transformation Hints

Authors: Xinrui Jiang, Heng Yu

Organizations: Department of Electrical Engineering Stanford University · Department of Computer Science Stanford University

Abstract

Creating sound effects for a new game-character skin requires a distinct acoustic identity while preserving gameplay-event roles. The challenge is to complete a coherent set of related sounds whose required degrees of redesign differ. We formulate this task as completion conditioned on base-skin audio, completed target assets, and a textual design description. We develop a pipeline to collect, process, and align corresponding events across League of Legends skins. Building on Stable Audio 3's pretrained audio prior, we fine-tune a latent inpainting model to jointly complete missing events. A signed soft retention mask encodes available audio and an adjustable transformation hint for each missing event, specifying the requested balance between retention and redesign. Experiments on held-out skins show improved reconstruction over the evaluated general-purpose audio editors. Target-derived hints further improve paired similarity, with three-level hints retaining most of the benefit of continuous guidance.

Figures & tables

Explore similar work

CardsList
  1. Spot, Separate, and Enhance: Fully Generative Approach for Audio Mixing

    Sep 24, 2026Ilpo Viertola, Giulio Cengarle, Gouthaman KV +2Audio UnderstandingMix

  2. InstructFX2FX: A Multi-Turn Text-to-Effect System for Sequential Audio Effect Refinement

    Jun 20, 2026Song-Ze Yu, Milan Liessens Dujardin, Yuxuan Cai +4Audio EditingAudio Understanding

  3. SongCraft: Unified Song Generation and Editing with Reconstructive Learning

    Sep 14, 2026Haohe Liu, Varun Nagaraja, Gael Le Lan +5Song Generation