Existing singing pitch correction approaches rely on target melodies or accompaniment tracks, which may be unavailable in practice. We formulate reference-free singing pitch correction as a music-constrained sequence editing task that determines whether and how each note should be corrected from the input performance alone. A pretrained symbolic music encoder with lightweight singing-domain adapters produces contextual representations of the singing MIDI. Based on these representations, two lightweight correction heads jointly model correction necessity and signed pitch modification through a factorized pitch-editing distribution. An input-dependent tonal prior derived from the estimated key distribution then reranks the candidate offsets, favoring tonally compatible corrections without target melodies or ground-truth key annotations. Experiments on real paired amateur and professional singing recordings show that the method improves note-level pitch accuracy from 73.25 to 83.43, outperforming a context-based baseline by 3.99 percentage points while balancing error correction and preservation of correctly performed notes.
Figures & tables
Figure 1: Overview of the proposed reference-free singing pitch correction framework.
Method
Acc. ↑
Δ Acc. ↑
Rep. ↑
Harm ↓
Input
73.25
0.00
–
–
Key Snapping
71.76
-1.49
20.76
9.62
BERT-APC [ 5 ]
79.44
+6.19
26.18
1.11
Ours
83.43
+10.18
46.81
3.19
Table 1: Objective results on the held-out real test set (%).
Figure 2: Qualitative and validation analyses.
Method
Acc. ↑
Δ Acc. ↑
Rep. ↑
Harm ↓
w/o adapters
76.87
+3.62
34.10
7.52
Flat edit head
82.35
+9.10
39.32
1.93
w/o reranking
81.52
+8.27
35.76
1.77
Ours
83.43
+10.18
46.81
3.19
Table 2: Ablation results on the held-out real test set (%).
Method
Pitch Corr. ↑
Nat. ↑
Overall ↑
Input
3.10
3.97
3.25
Key Snapping
3.26
3.80
3.67
BERT-APC
3.46
3.85
3.47
Ours
3.59
3.93
3.94
Table 3: Subjective evaluation on a five-point scale.
Automatic pitch correction (APC) requires distinguishing the unintended intonation errors from expressive pitch variation. Existing systems either lack explicit harmonic modeling, as vocal-only methods do, or do not directly use the note-level polyphonic context. Therefore, we propose MIDIBack, a note-level APC framework that jointly models the vocal and accompaniment events in a shared OctupleMIDI sequence. We evaluate MIDIBack under 6 note corruption regimes, including global outshift, learned note-dependent detuning, uniform perturbations, and their combinations. The resulting model achieves 78.6% overall raw pitch accuracy (RPA), and 81.5% under combined global outshift and learned detuning. Removing the accompaniment conditioning reduces RPA from 81.5% to 35.8% in outshift, showing the effectiveness of accompaniment context. Case studies on accompaniment modulation further illustrate that vocal note predictions
Joaquim Cavalcante, Yicheng Gu, Adriel Trajano +2
Centro de Inform´atica, Federal University of Para´ıba, Jo˜ao Pessoa, Brazil · Acoustic Lab, Aalto University, Espoo, Finland · MIR Research, Moises AI, Jo˜ao Pessoa, Brazil
Regenerating singing voices with altered lyrics while preserving melody consistency remains challenging, as existing methods either offer limited controllability or require laborious manual alignment. We propose YingMusic-Singer, a fully diffusion-based model enabling melody-controllable singing voice synthesis with flexible lyric manipulation. The model takes three inputs: an optional timbre reference, a melody-providing singing clip, and modified lyrics, without manual alignment. Trained with curriculum learning and Group Relative Policy Optimization, YingMusic-Singer achieves stronger melody preservation and lyric adherence than Vevo2, the most comparable baseline supporting melody control without manual alignment. We also introduce LyricEditBench, the first benchmark for melody-preserving lyric modification evaluation. The code, weights, benchmark, and demos are publicly available at https://github.com/ASLP-lab/YingMusic-Singer-Plus.
Song generation and editing have mostly been treated as separate tasks. Existing editing methods often require noise injection and regeneration or curated paired training data. We propose a unified approach for song generation and editing based on reconstructive pretraining, in which a model is trained to reconstruct audio from varying numbers of interpretable conditions. With conditions such as text and lyrics, the model learns to generate diverse songs. With dense conditions specifying fine-grained music attributes, the model learns to reconstruct the target and enables editing by modifying any single attribute while keeping others fixed. This leads to SongCraft, a latent flow matching based model trained for both generation and fine-grained editing. To improve song generation quality, we further introduce word-level phoneme alignment that improves pronunciation learning and accelerates convergence, beat conditioning that improves general musicality, and representation alignment on VAE latent space that produces semantically meaningful latents for improved generation quality. Experiments show that SongCraft achieves the lowest word error rate among evaluated song generation baselines while maintaining competitive audio quality. We further show that a single model can support editing of lyrics, vocal melody, beats, and singer identity, and we also study the trade-off between reconstruction quality and editability.