YuE2: Unifying Symbolic and Audio Music Generation at Frontier Quality
Organizations: HKUST HKGAI M-A-P Tokenwave.AI New York University Stanford University MBZUAI ACE Studio
Abstract
Symbolic models make melody, harmony, rhythm, and form explicit but typically stop before a finished recording; audio models produce complete songs while leaving composition implicit. We introduce YuE2, which unifies symbolic and audio music generation at frontier quality through symbolic planning. A single AR-NAR Mixture-of-Transformers (MoT) first writes a readable score specifying melody and harmony, expands it into semantic music tokens, and realizes it as full-song audio. In comparisons using the same checkpoint, experts prefer symbolic planning for overall quality and musicality, with 49.3% of overall preferences versus 34.6% without planning. Experts also favor the unified model over a separate language model and diffusion Transformer. On WildSongBench, YuE2 scores 6.73 on SongBench Global Avg, exceeding all evaluated public baselines. Selecting from eight candidates (best-of-8), YuE2 reaches 6.96, the highest observed mean among all evaluated systems. Expert listening further establishes its competitiveness with proprietary song generators, favoring best-of-8 over Suno v4.5 and yielding nearly balanced preferences against Suno v5. To learn this generation process from recordings without aligned scores, we introduce MERT2 and SheetSage2 to supply semantic and symbolic supervision. MERT2 sets a new state of the art in music representation learning, surpassing previous best results on 14 of 15 MARBLE metrics; SheetSage2 leads 12 of 15 benchmark-metric pairs in our lead-sheet transcription comparison. The same checkpoint follows score edits while largely preserving unedited musical content and generates zero-shot covers without cover-specific training. Its readable score also enables agentic music editing, with external language models translating user feedback into revisions of the composition.
Figures & tables
| SongBench | SongEval | AudioBox | Text alignment | Lyrics | |||||
| Model | Mus. | Avg | Mus. | Avg | PQ | MuLan | AMCaps | Q3O | PER |
| Suno v5 [ Suno, 2025 ] | 5.9918 | 6.8721 | 4.3051 | 4.3579 | 8.1698 | 0.5428 | 0.4353 | 4.5907 | 0.0810 |
| Suno v4.5 [ Team Suno, 2025 ] | 5.8317 | 6.6995 | 4.3198 | 4.3666 | 8.2541 | 0.5022 | 0.3873 | 4.4149 | 0.0580 |
| Suno v5.5 [ Shulman, 2026 ] | 5.8087 | 6.7150 | 4.1497 | 4.2152 | 8.1955 | 0.5089 | 0.3917 | 4.5914 | 0.0596 |
| Suno v6 [ Suno, 2026 ] | 5.6558 | 6.5562 | 4.2635 | 4.3086 | 8.1296 | 0.4916 | 0.4305 | 4.6258 | 0.0758 |
| Suno v6 Wild [ Suno, 2026 ] | 5.5644 | 6.4195 | 4.1716 | 4.2199 | 8.1785 | 0.4999 | 0.4316 | 4.5898 | 0.0745 |
| SongBench | SongEval | AudioBox | Text alignment | Lyrics | |||||
| Model | Mus. | Avg | Mus. | Avg | PQ | MuLan | AMCaps | Q3O | PER |
| YuE1 [ Yuan et al., 2025 ] | 4.0847 | 4.9165 | 3.1524 | 3.2150 | 7.8683 | 0.2623 | 0.2882 | 3.7301 | 0.3638 |
| SongBloom [ Yang et al., 2025a ] | 3.4493 | 4.2350 | 3.2048 | 3.2051 | 8.1539 | 0.2697 | 0.1926 | 3.0287 | 0.1919 |
| LeVo 2 [ Lei et al., 2026 ] | 5.4590 | 6.3247 | 3.9819 | 4.0234 | 8.3966 | 0.3542 | 0.2680 | 3.9458 | 0.2612 |
| ACE-Step 1.5 [ Gong et al., 2026 ] | 5.1588 | 6.0118 | 3.8051 | 3.8465 | 8.0518 | 0.4372 | 0.3869 | 4.5809 | 0.0746 |
| HeartMuLa [ Yang et al., 2026 ] | 5.4963 | 6.2483 | 4.5329 | 4.5519 | 8.2933 | 0.3823 | 0.2786 | 3.4907 | 0.1071 |
| Melody | Chords | Rhythm | Key | Tempo | Form | ||||
|---|---|---|---|---|---|---|---|---|---|
| Audio | Seq. | Seq. | F1 | Exact | Wtd. | Log err. | Acc. | Content | Bound. |
| Corresponding | 0.9464 | 0.9246 | 0.9396 | 0.9307 | 0.9515 | 0.0142 | 0.9869 | 0.7799 | 0.8724 |
| Mismatched | 0.2160 | 0.2473 | 0.6119 | 0.3004 | 0.3629 | 0.1131 | 0.6806 | 0.1547 | 0.1454 |
| Without score | 0.2055 | 0.1897 | 0.5875 | 0.2234 | 0.2831 | 0.1468 | 0.5366 | 0.1236 | 0.1245 |
| Edit | Metric | Score |
|---|---|---|
| Melody | Pitch accuracy | 84.17 |
| Harmony | Chord agreement | 79.54 |
| Rhythm | Relative onset accuracy | 73.43 |
| Key | Weighted key score [ Yuan et al., 2023 ] | 90.58 |
| Tempo | Acc2 [ Schreiber et al., 2020 ] | 95.68 |
| CLEWS | Discogs-VINet | Alignment | Quality | |||||||||
| Method | mAP | MRR | Hit@1 | Hit@5 | mAP | MRR | Hit@1 | Hit@5 | MuLan | Q3O | PQ | Mus. |
| SongEcho | 0.419 | 0.536 | 48.4 | 58.8 | 0.122 | 0.227 | 16.6 | 28.4 | 0.366 | 4.474 | 6.862 | 3.286 |
| ACE-Step 1.5 | 0.024 | 0.036 | 2.4 | 4.1 | 0.006 | 0.014 | 0.6 | 1.4 | 0.166 | 4.190 | 6.918 | 3.689 |
| YuE2 (full score) | 0.647 | 0.748 | 71.3 | 78.6 | 0.288 | 0.438 | 37.5 | 50.3 | 0.382 | 4.273 | 8.044 | 5.104 |
| Without chords | 0.598 | 0.715 | 67.3 | 76.1 | 0.179 | 0.304 | 23.6 | 36.6 | 0.417 | 4.482 | 8.117 | 5.490 |
| Without score | 0.006 | 0.008 | 0.3 | 0.9 | 0.004 | 0.008 | 0.2 | 0.8 | 0.474 | 4.837 | 8.186 | 5.691 |
| Dataset MTT GiantSteps GTZAN EmoMusic Task Tagging Key Genre Beat Emotion Model # Params ROC AP Acc. Refined Acc. F1 beat MERT-Large [ Li et al., 2024 ] 330M 90.6 37.9 64.1 77.6 86.8 56.7 76.1 Dasheng-1.2B [ Dinkel et al., 2024 ] 1.2B 91.5 40.4 58.0 81.4 87.7 57.4 75.0 MuQ [ Zhu et al., 2025 ] 310M 90.5 38.5 63.2 83.8 90.1 58.3 76.4 MusicFM [ Won et al., 2024 ] 330M 90.9 38.3 63.0 84.1 90.2 57.2 74.4 AudioMAE++ [ Yadav et al., 2025 ] 307M 91.2 39.5 61.7 80.3 90.0 59.0 75.7 MATPAC++ [ Quelennec et al., 2025 ] 307M 90.6 38.2 63.7 81.4 90.1 57.8 74.7 M2D-Large [ Gu et al., 2026 ] 307M 91.0 39.2 65.0 83.8 90.0 57.4 74.8 PupuJEPA-Large [ Gu et al., 2026 ] 307M 91.7 40.8 66.1 86.9 91.0 62.5 76.8 PupuJEPA-Huge [ Gu et al., 2026 ] 632M 91.3 39.7 64.8 85.9 90.5 62.0 78.5 MERT2-30s 632M 91.91 41.29 66.97 91.72 90.59 63.23 80.01 MERT2-FS 632M 91.74 41.20 67.05 90.69 90.57 63.52 78.14 |
| Dataset MTG-Jamendo Task Instrument Mood / theme Genre Top 50 Model # Params ROC AP ROC AP ROC AP ROC AP MERT-Large [ Li et al., 2024 ] 330M 75.5 18.8 75.3 13.5 86.1 18.0 82.6 29.1 Dasheng-1.2B [ Dinkel et al., 2024 ] 1.2B 75.0 19.0 76.1 15.5 85.5 18.8 82.4 29.6 MuQ [ Zhu et al., 2025 ] 310M 74.8 19.1 73.7 13.2 85.4 19.1 83.0 30.2 MusicFM [ Won et al., 2024 ] 330M 74.6 18.5 74.9 14.1 85.3 19.4 81.9 29.7 AudioMAE++ [ Yadav et al., 2025 ] 307M 77.1 19.9 75.6 14.0 86.3 18.9 83.1 31.1 MATPAC++ [ Quelennec et al., 2025 ] 307M 77.2 19.7 75.1 14.1 85.7 19.6 82.5 30.2 M2D-Large [ Gu et al., 2026 ] 307M 76.6 19.3 74.6 14.3 85.5 19.2 82.5 29.6 PupuJEPA-Large [ Gu et al., 2026 ] 307M 78.4 21.2 76.2 15.3 86.1 20.1 82.8 30.5 PupuJEPA-Huge [ Gu et al., 2026 ] 632M 77.6 20.5 75.9 14.7 85.9 20.1 83.1 30.7 MERT2-30s 632M 80.27 22.89 79.44 16.68 88.01 21.22 84.18 32.17 MERT2-FS 632M 80.27 23.51 78.74 15.74 87.98 20.66 84.13 31.62 |
| Token stream | Genre | Emotion | Key | MTT | Beat | ||||
|---|---|---|---|---|---|---|---|---|---|
| Representation | Rate (Hz) | Dim. | kb/s | Acc. | Score | AUROC | AP | F1 | |
| MERT2 tokenizer pre-VQ | 25 | 1024 | – | 84.14 | 66.19 | 56.94 | 91.60 | 40.02 | 89.81 |
| MERT2 tokenizer post-VQ | 25 | 1024 | 0.375 | 65.86 | 54.49 | 52.20 | 90.06 | 35.65 | 86.38 |
| LeVo 2 post-VQ | 25 | 1024 | 0.350 | 44.83 | 30.17 | 61.23 | 86.72 | 30.79 | 79.72 |
| EnCodec post-VQ | 75 | 128 | 1.500 | 31.72 | 26.51 | 15.76 | 80.84 | 22.32 | 76.50 |
| Task | Benchmark | Metric | SheetSage1 | Madmom | Prior specialist | SheetSage2-AR |
| Beat | GTZAN | F1 | 86.07 | 86.07 | 89.01 [ Foscarin et al., 2024 ] | 86.27 |
| osu2017 | 91.80 | 91.80 | 89.19 [ Foscarin et al., 2024 ] | 93.01 | ||
| Downbeat | GTZAN | F1 | 64.65 | 64.65 | 78.28 [ Foscarin et al., 2024 ] | 80.45 |
| osu2017 | 83.47 | 83.47 | 85.90 [ Foscarin et al., 2024 ] | 92.90 | ||
| Key | GiantSteps | Score | 43.89 | 74.62 | 72.09 [ Kong et al., 2025 ] | 77.73 |
| GTZAN | 54.56 | 72.05 | 74.43 [ Kong et al., 2025 ] | 75.77 |
Appendix figures & tables16 assets
Supplementary material from the paper’s appendix.
Appendix
| MERT2-30s | MERT2-FS | |||
|---|---|---|---|---|
| Task | Representation | Learning rate | Representation | Learning rate |
| GTZAN genre | L23 | L24 | ||
| GTZAN beat | L21 | L23 | ||
| GiantSteps key | L4 | L23 | ||
| EmoMusic | All | L24 | ||
| MTT | L22 | L23 | ||
| <time_8.48s> <beat_4_of_4> // beat <pitch_80_track_0> <duration_1.0> <subbeat_shift_1.0> <time_9.00s> <beat_1_of_4> // downbeat <structure_chorus> <chord_F#:maj> <pitch_82_track_0> <duration_0.5> <subbeat_shift_0.5> <pitch_73_track_0> <duration_0.5> <subbeat_shift_0.5> <time_9.52s> <beat_2_of_4> // beat <pitch_78_track_0> <duration_0.5>… |
| Task | Benchmark | Metric | SheetSage2-Prober | SheetSage2-AR |
| Beat | GTZAN | F1 | 82.93 | 86.27 |
| osu2017 | F1 | 92.28 | 93.01 | |
| Downbeat | GTZAN | F1 | 78.74 | 80.45 |
| osu2017 | F1 | 92.79 | 92.90 | |
| Key | GiantSteps | Score | 78.29 | 77.73 |
| GTZAN | Score | 72.62 | 75.77 |
| Model | Audio hours | Training stage |
|---|---|---|
| MERT2 | 700,000 | Foundation Pretraining |
| SheetSage2-AR | 28,400 | Autoregressive transcription |
| YuE2 | 346,000 | Joint symbolic and audio generation |
| Property | Value |
|---|---|
| Parameters | approximately 3.58B |
| Layers / hidden width / feed-forward width | 28 / 2,048 / 6,144 |
| Query heads / key–value groups / head width | 16 / 8 / 128 |
| Context length | 24,576 positions |
| Text vocabulary | 151,643 entries |
| Semantic vocabulary | 32,768 entries |
| Training task | AR prediction targets | Acoustic conditioning |
|---|---|---|
| Full generation | Score, then semantic tokens | Text, lyrics, score, and semantic tokens |
| No symbolic planning | Semantic tokens | Text, lyrics, and semantic tokens |
| No semantic tokens | Score | Text, lyrics, and score |
| Direct audio generation | None | Text and lyrics |
| Phase | Updates used | Max. context (positions) | Learning-rate schedule |
|---|---|---|---|
| AR pretraining | 60,000 | 16,384 | 5,000-step linear warmup to ; cosine decay with a floor at step 95,000. |
| Joint AR–NAR training | 40,000 | 24,576 | 5,000-step linear warmup to ; cosine decay with a floor at step 150,000. |
| Annealing | 6,000 | 24,576 | No warmup; cosine decay from approximately to over 6,000 updates. |
| Operation | Symbolic score | Generated tokens and latents |
|---|---|---|
| Create | ||
| Edit | user-modified | |
| Cover | transcribed from reference |
| Measure | Corresponding | Mismatched | Without score |
|---|---|---|---|
| Melody sequence | 0.9465 | 0.3152 | 0.3050 |
| Chord sequence | 0.9245 | 0.4658 | 0.4199 |
| Key: exact agreement | 0.9280 | 0.7527 | 0.6429 |
| Key: weighted agreement | 0.9489 | 0.8157 | 0.7320 |
| Musical form | 0.7798 | 0.2749 | 0.2276 |
| Tempo: half/double tolerant | 1.0000 | 0.7094 | 0.5654 |
| Condition | Onset accuracy | Coverage | Conditional accuracy |
|---|---|---|---|
| Unedited | 0.00 | 88.05 | 0.00 |
| Edited | 73.43 | 80.36 | 91.37 |
| MIDI: target exchange | 94.08 | 94.08 | 100.00 |
| MIDI: unchanged | 0.00 | 94.08 | 0.00 |
| MIDI: opposite shift | 0.00 | 93.54 | 0.00 |
| Condition | Quality | PER | CER |
|---|---|---|---|
| Unedited | 6.704 | 0.184 | 0.181 |
| Melody | 6.717 | 0.191 | 0.181 |
| Harmony | 6.674 | 0.179 | 0.175 |
| Key | 6.639 | 0.190 | 0.188 |
| Unedited (rhythm) | 6.728 | 0.214 | 0.197 |
| Rhythm | 6.731 | 0.185 | 0.175 |
| Criterion | Definition |
|---|---|
| Overall quality | Overall preference for the complete song, considering its composition, performance, sound quality, and fit to the request. |
| Musicality | Musical appeal and expressive coherence, including memorable melodic phrases, convincing harmony, rhythmic flow, effective arrangement, and development across sections. |
| Text alignment | Adherence to the supplied style description and lyrics, including the requested genre, mood, instruments, vocal configuration, tempo, and accurate delivery of the words. |
| Audio quality | Fidelity of the complete recording: audible detail, balanced levels, separation of sound sources, and freedom from unintended noise, clipping, distortion, or synthesis artifacts. |
| Vocals | Quality of the singing: natural vocal tone, stable pitch and timing, intelligible pronunciation, and expressive phrasing. |
| Accompaniment | Quality of the instrumental performance: believable timbres, clear articulation, expressive dynamics, and coordination among instruments and with the voice. |