Symbolic models make melody, harmony, rhythm, and form explicit but typically stop before a finished recording; audio models produce complete songs while leaving composition implicit. We introduce YuE2, which unifies symbolic and audio music generation at frontier quality through symbolic planning. A single AR-NAR Mixture-of-Transformers (MoT) first writes a readable score specifying melody and harmony, expands it into semantic music tokens, and realizes it as full-song audio. In comparisons using the same checkpoint, experts prefer symbolic planning for overall quality and musicality, with 49.3% of overall preferences versus 34.6% without planning. Experts also favor the unified model over a separate language model and diffusion Transformer. On WildSongBench, YuE2 scores 6.73 on SongBench Global Avg, exceeding all evaluated public baselines. Selecting from eight candidates (best-of-8), YuE2 reaches 6.96, the highest observed mean among all evaluated systems. Expert listening further establishes its competitiveness with proprietary song generators, favoring best-of-8 over Suno v4.5 and yielding nearly balanced preferences against Suno v5. To learn this generation process from recordings without aligned scores, we introduce MERT2 and SheetSage2 to supply semantic and symbolic supervision. MERT2 sets a new state of the art in music representation learning, surpassing previous best results on 14 of 15 MARBLE metrics; SheetSage2 leads 12 of 15 benchmark-metric pairs in our lead-sheet transcription comparison. The same checkpoint follows score edits while largely preserving unedited musical content and generates zero-shot covers without cover-specific training. Its readable score also enables agentic music editing, with external language models translating user feedback into revisions of the composition.
Figures & tables
Figure 1: YuE2 achieves frontier song quality, competitive with the evaluated proprietary systems. Automatic evaluation on WSB (192 prompts), including Suno v6 and v6 Wild [ Suno, 2026 ] . The quality index combines standardized SongBench [ Wu et al., 2026 ] and SongEval [ Yao et al., 2025 ] Global Avg; alignment combines standardized MuQ-MuLan [ Zhu et al., 2025 ] , AllMusicCaps [ Alonso-Jiménez et al., 2026 ] , and our Qwen3-Omni [ Xu et al., 2025 ] score. Bubble area: AudioBox production quality (PQ) [ Tjandra et al., 2025 ] ; black outlines: the observed Pareto front on these two indices. Both YuE2 settings use symbolic planning; Bo8 means best-of-8. Appendix A.4 gives index construction and weight sensitivity; Figure 7 reports expert preferences.
Figure 2: YuE2 progressively commits to musical detail while keeping the symbolic score explicit and inspectable.
Figure 3: Composing in symbols, performing in audio. YuE2 generates an editable symbolic composition s , 25-Hz MERT2 semantic tokens c , and 25-Hz continuous acoustic latents z before 48-kHz stereo variational-autoencoder (VAE) decoding [ Kingma & Welling, 2013 ] . Creation: YuE2 generates s from the conditioning text. Covering: SheetSage2 extracts s from a reference recording. Editing: the user supplies a modified s . The same YuE2 checkpoint then generates the semantic and acoustic representations in all three modes. Within each of 28 layers, side-by-side AR and NAR experts use separate normalization, projections, and MLPs around one shared attention computation. The AR stream conditions on style and lyric text y and performs causal next-token prediction over s and c ; the NAR stream predicts flow velocity over zt , where t∈[0,1] is flow time, with bidirectional latent attention.
Figure 4: MERT2 learns from discrete targets jointly derived from two encoder views of the same audio. (A) Multi-View Target Synthesis. Frozen MuQ [ Zhu et al., 2025 ] and Qwen2-Audio-Instruct [ Chu et al., 2024 ] encoders map the same recording to time-aligned features. A shared four-level residual vector quantizer produces four cached 25-Hz code streams. Separate decoders reconstruct both feature spaces from the same quantized representation. (B) Stage 1 — Foundation Pretraining. A bidirectional ConvNeXt–Conformer encoder [ Liu et al., 2022 ; Gulati et al., 2020 ] is trained to predict these codes from masked audio. (C) Adaptation. Stage 2-FS — Full-Song Adaptation continues pretraining on complete recordings lasting 30–360 seconds while retaining bidirectional context, producing MERT2-FS for SheetSage2. Stage 2 — Causal Adaptation switches self-attention to a causal mask and continues along the semantic-tokenizer branch.
Figure 5: SheetSage2-AR architecture and output representation. A full-context MERT2-FS encoder with a frozen backbone and trainable low-rank adapters supplies a learned layer mixture to a six-layer autoregressive decoder. Task prompts select the target attributes, and grammar-constrained decoding produces a chronological sequence of timing, metrical, structural, harmonic, and melodic events. A deterministic builder converts the event sequence to ABC notation, which is rendered as a lead sheet. The right panel uses a reference bar to illustrate this representation. The excerpt is serialized independently, with the active key emitted at its first beat and omitted thereafter unless it changes.
Figure 6: MERT2 tokenizer training and deployment. (A) Causal Adaptation initializes from Stage 1, masks temporal spans of the input audio, switches the 24-layer Conformer [ Gulati et al., 2020 ] to causal attention, and continues the Stage 1 objective: predicting all four cached target streams at masked positions under LMLM . (B) Supervised Fine-Tuning continues the adapted encoder on full songs and adds lyric connectionist temporal classification (CTC) [ Graves et al., 2006 ] plus mel and chroma reconstruction. (C) Semantic Quantization inserts a learned 32-dimensional, 32,768-entry clustered vector quantizer (CVQ) between layers 13 and 14. Adapted from CVQ-VAE [ Zheng & Vedaldi, 2023 ] , its online feature-anchor update returns rarely selected entries to the assignment distribution. During Stage 4, layers 14–23 and the CTC, mel, and chroma heads remain above the quantizer, allowing their losses to supervise the quantized states. (D) Deployment retains the log-mel/ConvNeXt frontend [ Liu et al., 2022 ] , Conformer layers 0–13 with causal self-attention, and the quantizer, producing one semantic ID c every 40 ms (25 Hz; 375 bit/s) for autoregressive modeling and acoustic conditioning in YuE2.
SongBench
SongEval
AudioBox
Text alignment
Lyrics
Model
Mus. ↑
Avg ↑
Mus. ↑
Avg ↑
PQ ↑
MuLan ↑
AMCaps ↑
Q3O ↑
PER ↓
Suno v5 [ Suno, 2025 ]
5.9918
6.8721
4.3051
4.3579
8.1698
0.5428
0.4353
4.5907
0.0810
Suno v4.5 [ Team Suno, 2025 ]
5.8317
6.6995
4.3198
4.3666
8.2541
0.5022
0.3873
4.4149
0.0580
Suno v5.5 [ Shulman, 2026 ]
5.8087
6.7150
4.1497
4.2152
8.1955
0.5089
0.3917
4.5914
0.0596
Suno v6 [ Suno, 2026 ]
5.6558
6.5562
4.2635
4.3086
8.1296
0.4916
0.4305
4.6258
0.0758
Suno v6 Wild [ Suno, 2026 ]
5.5644
6.4195
4.1716
4.2199
8.1785
0.4999
0.4316
4.5898
0.0745
Table 1: Comparison with proprietary systems on WildSongBench (WSB). All metrics cover 192 prompts. Mus.: Musicality; Avg: mean of all SongBench [ Wu et al., 2026 ] or SongEval [ Yao et al., 2025 ] dimensions; PQ: AudioBox production quality [ Tjandra et al., 2025 ] ; MuLan: MuQ-MuLan [ Zhu et al., 2025 ] ; AMCaps: AllMusicCaps [ Alonso-Jiménez et al., 2026 ] ; Q3O: Qwen3-Omni prompt adherence [ Xu et al., 2025 ] ; PER: phoneme error rate. Both YuE2 settings use symbolic planning. The standard comparison selects the lower-PER output from two candidates per prompt for every system; YuE2 best-of-8 (Bo8) selects from eight, prioritizing SongBench Musicality, then Q3O, then PER. Appendix A.3 details selection. MM abbreviates MiniMax. Bold and underlining mark the best and second-best observed means; higher is better except PER.
SongBench
SongEval
AudioBox
Text alignment
Lyrics
Model
Mus. ↑
Avg ↑
Mus. ↑
Avg ↑
PQ ↑
MuLan ↑
AMCaps ↑
Q3O ↑
PER ↓
YuE1 [ Yuan et al., 2025 ]
4.0847
4.9165
3.1524
3.2150
7.8683
0.2623
0.2882
3.7301
0.3638
SongBloom [ Yang et al., 2025a ]
3.4493
4.2350
3.2048
3.2051
8.1539
0.2697
0.1926
3.0287
0.1919
LeVo 2 [ Lei et al., 2026 ]
5.4590
6.3247
3.9819
4.0234
8.3966
0.3542
0.2680
3.9458
0.2612
ACE-Step 1.5 [ Gong et al., 2026 ]
5.1588
6.0118
3.8051
3.8465
8.0518
0.4372
0.3869
4.5809
0.0746
HeartMuLa † [ Yang et al., 2026 ]
5.4963
6.2483
4.5329
4.5519
8.2933
0.3823
0.2786
3.4907
0.1071
Table 2: Comparison with publicly released systems on WSB. Prompts, metrics, abbreviations, and YuE2 selection budgets follow Table 1 . Both YuE2 rows use symbolic planning. HeartMuLa † reports using SongEval and AudioBox during post-training to filter supervised fine-tuning data and construct DPO preference pairs [ Yang et al., 2026 ] . Bold and underlined values indicate the best and second-best scores across all rows, respectively. Appendix A.3 gives the baseline protocols.
Figure 7: Expert preferences against proprietary song generators. Columns show YuE2 and YuE2 (best-of-8), both with symbolic planning. Dark blue favors the YuE2 setting named above the column; light blue favors the baseline named in the row. Baselines are Suno v4.5 [ Team Suno, 2025 ] , v5 [ Suno, 2025 ] , v5.5 [ Shulman, 2026 ] , v6 and v6 Wild [ Suno, 2026 ] , and Mureka 9 [ Mureka, 2026 ] .
Figure 8: Symbolic planning improves perceived song quality, melody, and chord progression. Expert preferences for generation with versus without symbolic planning. Dark blue favors planning, medium blue denotes ties, and light blue favors generation without planning. Appendix D.4 details the listening protocol.
Figure 9: Planning develops a recurring melodic idea in this paired example. Both songs use the same prompt and lyrics. Read each staff from left to right; higher notes indicate higher pitch, and letters above the staff name the accompanying chords. Purple marks phrases with repeated lyric openings. With planning (left), the verse varies a low–low–high–high note pattern; the chorus recalls it an octave higher with longer phrase endings, linking the sections through recognizable variation. Without planning (right), the verse repeats the same B – D – E – E pattern, with less variation and a less direct melodic connection to the chorus. Gray marks a chord transition where experts hear the bass descent break, weakening harmonic continuity. The left score starts from the generated plan and the right from a SheetSage2 transcription; experts corrected both against the completed recordings for display, without regenerating audio. Appendix gives the full scores and further analysis.
Figure 10: Experts prefer unified generation on five of six criteria; text alignment is nearly balanced. YuE2’s Mixture-of-Transformers (MoT) [ Deng et al., 2025 ] learns semantic-token prediction and acoustic generation in one backbone; the baseline uses a separate language model and diffusion Transformer (LM+DiT) [ Xu et al., 2026 ] . Both use melody-and-chord planning, the same training-data volume and representations, and the same prompts and two-candidate selection protocol. Dark blue favors YuE2, medium blue denotes ties, and light blue favors LM+DiT. Bars show percentages of judged expert responses (206–208 per criterion), excluding unable-to-judge answers.
Melody
Chords
Rhythm
Key
Tempo
Form
Audio
Seq. ↑
Seq. ↑
F1 ↑
Exact ↑
Wtd. ↑
Log err. ↓
Acc. ↑
Content ↑
Bound. ↑
Corresponding
0.9464
0.9246
0.9396
0.9307
0.9515
0.0142
0.9869
0.7799
0.8724
Mismatched
0.2160
0.2473
0.6119
0.3004
0.3629
0.1131
0.6806
0.1547
0.1454
Without score
0.2055
0.1897
0.5875
0.2234
0.2831
0.1468
0.5366
0.1236
0.1245
Table 3: Score–audio consistency on WSB. Mismatched audio uses the other score for the same prompt. Rhythm reports vocal onset-interval F1. Values are prompt means. Tempo accuracy allows 8% error; lower log error is better.
Edit
Metric
Score ↑
Melody
Pitch accuracy
84.17
Harmony
Chord agreement
79.54
Rhythm
Relative onset accuracy
73.43
Key
Weighted key score [ Yuan et al., 2023 ]
90.58
Tempo
Acc2 [ Schreiber et al., 2020 ]
95.68
Table 4: Score editing on WildSongBench. Higher is better for all scores (0–100).
CLEWS
Discogs-VINet
Alignment
Quality
Method
mAP ↑
MRR ↑
Hit@1 ↑
Hit@5 ↑
mAP ↑
MRR ↑
Hit@1 ↑
Hit@5 ↑
MuLan ↑
Q3O ↑
PQ ↑
Mus. ↑
SongEcho
0.419
0.536
48.4
58.8
0.122
0.227
16.6
28.4
0.366
4.474
6.862
3.286
ACE-Step 1.5
0.024
0.036
2.4
4.1
0.006
0.014
0.6
1.4
0.166
4.190
6.918
3.689
YuE2 (full score)
0.647
0.748
71.3
78.6
0.288
0.438
37.5
50.3
0.382
4.273
8.044
5.104
Without chords
0.598
0.715
67.3
76.1
0.179
0.304
23.6
36.6
0.417
4.482
8.117
5.490
Without score
0.006
0.008
0.3
0.9
0.004
0.008
0.2
0.8
0.474
4.837
8.186
5.691
Table 5: YuE2 with full scores leads both evaluated cover systems on all eight retrieval measures without cover-specific training. SHS100K [ NovaFrost, 2018 ] evaluation across 948 works, with 3,792 outputs per method and no candidate selection. These works are absent from YuE2’s training corpus. Each query ranks 10,545 recordings after excluding its source. All metrics are higher-is-better.
Table 6: MERT2-30s leads on 14 of 15 MARBLE [ Yuan et al., 2023 ] metrics against the listed external baselines. Baseline scores are from Tables III and V of Gu et al. [2026] (arXiv v3). AudioMAE++, MATPAC++, and M2D-Large are that study’s retrained models; the remaining external baselines use its evaluations of released checkpoints or its own PupuJEPA models. Results cover MagnaTagATune (MTT) [ Law et al., 2009 ] , GiantSteps [ Knees et al., 2015 ] , GTZAN [ Tzanetakis & Cook, 2002 ] , EmoMusic [ Soleymani et al., 2013 ] , and MTG-Jamendo [ Bogdanov et al., 2019 ] . MERT2-FS (full-song) continues pretraining on complete recordings lasting 30–360 seconds.
Token stream
Genre
Emotion
Key
MTT
Beat
Representation
Rate (Hz)
Dim.
kb/s
Acc. ↑
R2↑
Score ↑
AUROC ↑
AP ↑
F1 ↑
MERT2 tokenizer pre-VQ
25
1024
–
84.14
66.19
56.94
91.60
40.02
89.81
MERT2 tokenizer post-VQ
25
1024
0.375
65.86
54.49
52.20
90.06
35.65
86.38
LeVo 2 post-VQ
25
1024
0.350
44.83
30.17
61.23
86.72
30.79
79.72
EnCodec post-VQ
75
128
1.500
31.72
26.51
15.76
80.84
22.32
76.50
Table 7: At 375 bits/s, the MERT2 tokenizer leads five of six metrics among the evaluated quantized representations. Scores are multiplied by 100 (higher is better), with the best post-VQ result per column in bold. Emotion reports the mean R2 across arousal and valence. Rate, feature dimension, and nominal bitrate characterize each representation; external controls are LeVo 2 [ Lei et al., 2026 ] and EnCodec [ Défossez et al., 2022 ] .
Task
Benchmark
Metric
SheetSage1
Madmom
Prior specialist
SheetSage2-AR
Beat
GTZAN
F1 ↑
86.07
86.07
89.01 [ Foscarin et al., 2024 ]
86.27
osu2017
91.80
91.80
89.19 [ Foscarin et al., 2024 ]
93.01
Downbeat
GTZAN
F1 ↑
64.65
64.65
78.28 [ Foscarin et al., 2024 ]
80.45
osu2017
83.47
83.47
85.90 [ Foscarin et al., 2024 ]
92.90
Key
GiantSteps
Score ↑
43.89
74.62
72.09 [ Kong et al., 2025 ]
77.73
GTZAN
54.56
72.05
74.43 [ Kong et al., 2025 ]
75.77
Table 8: SheetSage2-AR: one model for full-song lead-sheet analysis. The comparison includes SheetSage1 [ Donahue et al., 2022 ] , madmom [ Böck et al., 2016 ] , and task-specific systems cited beside their scores. SheetSage2-AR achieves the highest scores on 12 of 15 benchmark–metric pairs. Scores are percentages; higher is better.
Appendix figures & tables16 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 11: YuE2 achieves the highest Musicality index among the evaluated public generators, with fewer parameters than six of eight baselines. Results use 192 WildSongBench prompts. The index combines standardized SongBench and SongEval Musicality scores with 2:1 weights; it is not a percentage. Moving right indicates higher Musicality index; moving up indicates fewer parameters on a reversed logarithmic axis. Circles denote the two-candidate settings; the hollow diamond denotes YuE2 (best-of-8). The connecting segment identifies the same 3.58B-parameter model under these two selection protocols. Counts include generation networks and conditioning encoders, excluding audio tokenizers, VAEs, and vocoders. Appendix A.5 gives the index and counting details.
MERT2-30s
MERT2-FS
Task
Representation
Learning rate
Representation
Learning rate
GTZAN genre
L23
5×10−3
L24
5×10−4
GTZAN beat
L21
1×10−3
L23
1×10−3
GiantSteps key
L4
1×10−3
L23
1×10−3
EmoMusic
All
5×10−5
L24
5×10−4
MTT
L22
1×10−3
L23
1×10−3
Appendix
Table 9: MARBLE probe hyperparameters for the reported MERT2 results. Lk denotes the output of Conformer layer k (numbered 1–24). “All” concatenates the 24 layer representations and applies a learned linear projection to 1,024 dimensions.
Table 10: Example event sequence decoded by SheetSage2-AR.
Task
Benchmark
Metric
SheetSage2-Prober
SheetSage2-AR
Beat
GTZAN
F1 ↑
82.93
86.27
osu2017
F1 ↑
92.28
93.01
Downbeat
GTZAN
F1 ↑
78.74
80.45
osu2017
F1 ↑
92.79
92.90
Key
GiantSteps
Score ↑
78.29
77.73
GTZAN
Score ↑
72.62
75.77
Appendix
Table 11: SheetSage2-AR compared with SheetSage2-Prober. The models use the same benchmark cohorts and scoring protocol as Table 8 . All scores are percentages; bold marks the higher value in each row.
Model
Audio hours
Training stage
MERT2
700,000
Foundation Pretraining
SheetSage2-AR
28,400
Autoregressive transcription
YuE2
346,000
Joint symbolic and audio generation
Appendix
Table 12: Training corpus sizes in rounded audio hours.
Property
Value
Parameters
approximately 3.58B
Layers / hidden width / feed-forward width
28 / 2,048 / 6,144
Query heads / key–value groups / head width
16 / 8 / 128
Context length
24,576 positions
Text vocabulary
151,643 entries
Semantic vocabulary
32,768 entries
Appendix
Table 13: Main YuE2 architecture.
Training task
AR prediction targets
Acoustic conditioning
Full generation
Score, then semantic tokens
Text, lyrics, score, and semantic tokens
No symbolic planning
Semantic tokens
Text, lyrics, and semantic tokens
No semantic tokens
Score
Text, lyrics, and score
Direct audio generation
None
Text and lyrics
Appendix
Table 14: Four training tasks mixed within one YuE2 model. The four tasks are sampled in a balanced ratio during joint training and annealing.
Phase
Updates used
Max. context (positions)
Learning-rate schedule
AR pretraining
60,000
16,384
5,000-step linear warmup to 10−4 ; cosine decay with a 10−5 floor at step 95,000.
Joint AR–NAR training
40,000
24,576
5,000-step linear warmup to 3×10−4 ; cosine decay with a 10−4 floor at step 150,000.
Annealing
6,000
24,576
No warmup; cosine decay from approximately 2.726×10−4 to 3×10−5 over 6,000 updates.
Appendix
Table 15: Training phases and optimization schedules for YuE2.
Operation
Symbolic score
Generated tokens and latents
Create
s∼p(s∣y)
c,z∼p(c,z∣y,s)
Edit
user-modified s~
c~,z~∼p(c,z∣y,s~)
Cover
s^ transcribed from reference
c′,z′∼p(c,z∣y′,s^)
Appendix
Table 16: Score conditioning for creation, editing, and cover generation.
Measure
Corresponding
Mismatched
Without score
Melody sequence ↑
0.9465
0.3152
0.3050
Chord sequence ↑
0.9245
0.4658
0.4199
Key: exact agreement ↑
0.9280
0.7527
0.6429
Key: weighted agreement ↑
0.9489
0.8157
0.7320
Musical form ↑
0.7798
0.2749
0.2276
Tempo: half/double tolerant ↑
1.0000
0.7094
0.5654
Appendix
Table 17: Companion comparisons allowing global transposition or tempo ambiguity. The first five rows share one melody/chord-derived pitch shift per comparison. The last row permits half or double the target tempo, with the same 8% tolerance. Higher is better. Values are prompt means.
Condition
Onset accuracy
Coverage
Conditional accuracy
Unedited
0.00
88.05
0.00
Edited
73.43
80.36
91.37
MIDI: target exchange
94.08
94.08
100.00
MIDI: unchanged
0.00
94.08
0.00
MIDI: opposite shift
0.00
93.54
0.00
Appendix
Table 18: Rhythm exchange accuracy and correspondence coverage. All scores are percentages.
Condition
Quality ↑
PER ↓
CER ↓
Unedited
6.704
0.184
0.181
Melody
6.717
0.191
0.181
Harmony
6.674
0.179
0.175
Key
6.639
0.190
0.188
Unedited (rhythm)
6.728
0.214
0.197
Rhythm
6.731
0.185
0.175
Appendix
Table 19: Song quality and lyric retention after editing on WildSongBench.
Criterion
Definition
Overall quality
Overall preference for the complete song, considering its composition, performance, sound quality, and fit to the request.
Musicality
Musical appeal and expressive coherence, including memorable melodic phrases, convincing harmony, rhythmic flow, effective arrangement, and development across sections.
Text alignment
Adherence to the supplied style description and lyrics, including the requested genre, mood, instruments, vocal configuration, tempo, and accurate delivery of the words.
Audio quality
Fidelity of the complete recording: audible detail, balanced levels, separation of sound sources, and freedom from unintended noise, clipping, distortion, or synthesis artifacts.
Vocals
Quality of the singing: natural vocal tone, stable pitch and timing, intelligible pronunciation, and expressive phrasing.
Accompaniment
Quality of the instrumental performance: believable timbres, clear articulation, expressive dynamics, and coordination among instruments and with the voice.
Appendix
Table 20: Definitions of the six expert-listening criteria.
Figure 12: Expert preferences for YuE2 versus six proprietary song generators across all six criteria. Audio-quality preferences favor YuE2 against all six baselines; overall preferences favor it in two comparisons. YuE2 uses melody-and-chord planning and selects the candidate with fewer lyric errors from two generated songs. Baselines are Suno v4.5 [ Team Suno, 2025 ] , v5 [ Suno, 2025 ] , v5.5 [ Shulman, 2026 ] , v6 and v6 Wild [ Suno, 2026 ] , and Mureka 9 [ Mureka, 2026 ] . Dark blue favors YuE2, light blue favors the named baseline, and medium blue denotes ties. Baseline order is shared by the left and right panels. Each bar sums to 100% of judged responses, excluding unable-to-judge answers. Counts and CR1 intervals are in Appendix D.4.4 .
Figure 13: Expert preferences for YuE2 (best-of-8) versus six proprietary song generators across all six criteria. Overall preferences favor best-of-8 over Suno v4.5, are nearly balanced against v5, and favor v6 over best-of-8. Audio-quality preferences favor best-of-8 in five of six comparisons. YuE2 uses melody-and-chord planning and selects from eight candidates by musicality, prompt adherence, then lyric accuracy. Baselines are Suno v4.5 [ Team Suno, 2025 ] , v5 [ Suno, 2025 ] , v5.5 [ Shulman, 2026 ] , v6 and v6 Wild [ Suno, 2026 ] , and Mureka 9 [ Mureka, 2026 ] . Dark blue favors YuE2 (best-of-8), light blue favors the named baseline, and medium blue denotes ties. Baseline order is shared by the left and right panels. Each bar sums to 100% of judged responses, excluding unable-to-judge answers. Counts and CR1 intervals are in Appendix D.4.4 .
Figure 14: Expert preferences across six criteria for symbolic-planning ablations and YuE2 versus separate LM+DiT. Full planning receives more preferences than no planning on every criterion; unified generation is preferred overall both with and without planning. Each panel contains five direct comparisons. Full planning includes melody and chords; melody-only planning omits chords. The first three rows compare planning conditions within YuE2. The last two compare YuE2’s AR–NAR Mixture-of-Transformers [ Deng et al., 2025 ] with a separate language model and diffusion Transformer (LM+DiT), with and without planning. Dark blue favors the first-named condition in each row; light blue favors the second. Medium blue denotes ties. Each bar sums to 100% of judged responses, excluding unable-to-judge answers. Counts and CR1 intervals are in Appendix D.4.4 .
We introduce Qwen-Music, a music generation model that produces high-fidelity songs with complete vocals. It supports text-to-music generation from descriptions, lyrics, and musical attributes, and cover song generation with different styles and vocal characteristics. Qwen-Music comprises three components: Qwen-Music-Tokenizer, Qwen-Music-LLM, and Qwen-Music-Render. The tokenizer compresses audio into a 25 Hz single-codebook stream of Music Semantic Tokens that preserve semantic and melodic information. The LLM performs autoregressive modeling with a melody-token-based chain-of-thought (Melody-CoT) mechanism that plans melodies before full-song generation, improving musicality, structural coherence, and reference-melody preservation. The renderer enriches discrete semantic tokens with acoustic details to produce high-fidelity stereo waveforms. We train the LLM using a quality-aware pre-training curriculum followed by progressive post-training with supervised initialization, offline DPO, and online GSPO to improve musicality and instruction following. On 600 evaluation inputs, Qwen-Music achieves state-of-the-art results in 13 of 16 objective musicality and audio-quality metrics. Professional evaluators also prefer Qwen-Music over leading proprietary systems. For cover generation, Qwen-Music preserves reference melodies more accurately than Suno V5.5, Suno V5, and MiniMax Cover on the AI-generated reference set, and outperforms MiniMax Cover on most metrics on the real-world popular-song reference set.
Symbolic music evaluation for large language models remains fragmented across representations, datasets, and metrics. We introduce LilyBench, a LilyPond-based benchmark that jointly evaluates symbolic music generation and music understanding on the same family of open-weight LLMs. The benchmark includes a 200-prompt generation suite and ten understanding tasks adapted from ABC-Eval, covering syntax, metadata prediction, structural sequencing, and music recognition. Generation quality is evaluated using compile rate, MusPy descriptor distributions via Jensen-Shannon similarity, and LilyBERT-based Fréchet Music Distance (FMD). Experiments on four open-weight models show that executable LilyPond generation is achievable in zero-shot settings, while structural understanding tasks remain challenging despite strong performance on composer and genre recognition. Our experiments also reveal systematic disagreements between descriptor-based and embedding-based metrics, suggesting that symbolic music evaluation benefits from metric triangulation rather than single-score ranking. We release the benchmark, prompt bank, and evaluation code to support future research in symbolic music generation and understanding at https://github.com/CSCPadova/lilybench
Matteo Spanio, Mohammad Torabi, Andrea Poltronieri +1
Centro di Sonologia Computazionale, University of Padova, Padova, Italy · Music Technology Group, Universitat Pompeu Fabra, Barcelona, Spain
Large language models can produce superficially legal twelve-tone scores that collapse into degenerate textures. We introduce a neuro-symbolic harness that wraps a language-model proposer in a generate-verify-repair-trace loop with symbolic verification. The complete pipeline improves event-local consistency without claiming whole-piece legality. Across 40 controlled tasks and four paired models, constraint-checked delivery rises from 13.3% under raw generation to 48.1% with the harness; it abstains on the remaining 51.9% of runs. The pass rate of a narrower collision and serialisation-consistency check rises from 33.5% to 58.3%, while degeneracy remains near 0.05, including under adversarial prompting. A blinded evaluation by five experts also shows a descriptive aggregate preference for harness candidates over raw generation in adherence, perceived legality, coherence, and overall quality.
Congren Dai, Danni Zhao, Enyang Liu +7
Central Conservatory of Music · Tsinghua University