Adapting a Latent Audio Diffusion Model to Historical Guqin Recordings: A Listening-Driven Case Study
Organizations: Division of Speech, Music and Hearing (TMH) KTH Royal Institute of Technology, Stockholm, Sweden
Abstract
We report a small-data case study in adapting a pretrained latent audio diffusion model to the guqin, the seven-string Chinese zither, aiming at an "AI radio" that plays guqin-style music without end. From a library of historical recordings we curate 412 solo performances (42.5 h, 61 performers) and split them by composition. We fine-tune a rank-16 DoRA adapter on Stable Audio 3 Medium using its continuous latents, masking weighted towards continuation, captions that combine researched notes on each piece with mood tags and an automatically estimated pentatonic mode, and random-length crops, stopping when held-out loss stops improving. Seven blind listening studies by one expert listener guided every decision. The final adapter was rated highest for continuing unseen pieces (3.9/5, against 3.4 for the best earlier adapter) and 4.6/5 for generating from free-written scene descriptions. Negative results are equally informative: a from-scratch autoregressive model over the same latents produced no recognisable timbre, a signal-level friction-noise measure correlated with the listener's complaints in the wrong direction, and a pentatonic-fit measure tracked ratings overall but barely within a group of candidates. Chaining continuations for long playback exposed a silent tail on every generated clip and a loudness feedback loop, both with simple fixes. With one listener and at most ten clips per condition, no paired difference is statistically significant; we present an exploratory record of what helped, what did not, and why.
Figures & tables
| Study | Condition | Mean score | |
|---|---|---|---|
| R1 masking | released default mask (10/80/10) / continuation mask (10/30/60) / 412 recordings, continuation mask | 8 | 2.88 / 2.62 / 3.38 |
| R2 inference | adapter strength 1.0 / 0.7 / adapter on early steps only | 9–10 | 3.56 / 3.60 / 3.67 |
| base model with 50 steps / earlier 222-clip adapter | 10 | 2.70 / 3.10 | |
| R3 captions | continuation: no text in training / no prompt / matching / opposite tags | 9 | 3.89 / 3.56 / 4.00 / 4.00 |
| text only: adapter trained without / with captions | 5 | 2.00 / 3.80 | |
| R4 latent LM | from-scratch latent LM / SA3 adapter | 11 / 6 | 1.00 / 2.00 |